Tag Archives: politics

Preserving U.S. Government Websites and Data as the Obama Term Ends

Posted on December 15, 2016 by jefferson

Long before the 2016 Presidential election cycle librarians have understood this often-overlooked fact: vast amounts of government data and digital information are at risk of vanishing when a presidential term ends and administrations change. For example, 83% of .gov pdf’s disappeared between 2008 and 2012.

That is why the Internet Archive, along with partners from the Library of Congress, University of North Texas, George Washington University, Stanford University, California Digital Library, and other public and private libraries, are hard at work on the End of Term Web Archive, a wide-ranging effort to preserve the entirety of the federal government web presence, especially the .gov and .mil domains, along with federal websites on other domains and official government social media accounts.

While not the only project the Internet Archive is doing to preserve government websites, ftp sites, and databases at this time, the End of Term Web Archive is a far reaching one.

The Internet Archive is collecting webpages from over 6,000 government domains, over 200,000 hosts, and feeds from around 10,000 official federal social media accounts. The effort is likely to preserve hundreds of millions of individual government webpages and data and could end up totaling well over 100 terabytes of data of archived materials. Over its full history of web archiving, the Internet Archive has preserved over 3.5 billion URLs from the .gov domain including over 45 million PDFs.

This end-of-term collection builds on similar initiatives in 2008 and 2012 by original partners Internet Archive, Library of Congress, University of North Texas, and California Digital Library to document the “gov web,” which has no mandated, domain-wide single custodian. For instance, here is the National Institute of Literacy (NIFL) website in 2008. The domain went offline in 2011. Similarly, the Sustainable Development Indicators (SDI) site was later taken down. Other websites, such as invasivespecies.gov were later folded into larger agency domains. Every web page archived is accessible through the Wayback Machine and past and current End of Term specific collections are full-text searchable through the main End of Term portal. We have also worked with additional partners to provide access to the full data for use in data-mining research and projects.

The project has received considerable press attention this year, with related stories in The New York Times, Politico, The Washington Post, Library Journal, Motherboard, and others.

“No single government entity is responsible for archiving the entire federal government’s web presence,” explained Jefferson Bailey, the Internet Archive’s Director of Web Archiving. “Web data is already highly ephemeral and websites without a mandated custodian are even more imperiled. These sites include significant amounts of publicly-funded federal research, data, projects, and reporting that may only exist or be published on the web. This is tremendously important historical information. It also creates an amazing opportunity for libraries and archives to join forces and resources and collaborate to archive and provide permanent access to this material.”

This year has also seen a significant increase in citizen and librarian driven “hackathons” and “nomination-a-thons” where subject experts and concerned information professionals crowdsource lists of high-value or endangered websites for the End of Term archiving partners to crawl. Librarian groups in New York City are holding nomination events to make sure important sites are preserved. And universities such as The University of Toronto are holding events for “guerrilla archiving” focused specifically on preserving climate related data.

We need your help too! You can use the End of Term Nomination Tool to nominate any .gov or government website or social media site and it will be archived by the project team. If you have other ideas, please comment here or send ideas to info@archive.org. And you can also help by donating to the Internet Archive to help our continued mission to provide “Universal Access to All Knowledge.”

Pro-Airbnb advertising dominated recent political TV ads in San Francisco

Posted on November 4, 2015 by Nancy Watzman

Based on algorithmic analysis, Pro-Airbnb advertising dominated political TV ads in San Francisco in the weeks leading up to Election Day. Two thirds of the minutes devoted to political ads on several initiatives and races before voters focused on arguments against a proposal to curb the company’s operations in the city, according to a review of the Internet Archive television archive. Voters ended up rejecting Proposition F, whose opponents claimed it would encourage neighbors to spy on each other and increase lawsuits, by a margin of 55 to 45 percent.

The Archive identified total of 1,959 minutes of ads (4,591 plays) opposing Proposition F, out of 2,895 minutes devoted to all political TV ads, or roughly two thirds of the air-time.

To put that in perspective, Mayor Ed Lee, who won his reelection easily, was the subject of only 55 minutes of ads. Though he appeared in and narrated hundreds of ads supporting Propositions A and D, the only ads that mention his mayoral race were airings of a support ad paid for not by his own campaign, but rather by an independent expenditure from Clint Reilly, a local real estate developer and former professional political consultant.

Samples of all ads found to be related to 2015 San Francisco elections can be viewed here, and metadata about those that occurred in archived television can be downloaded from this page.

The only political ad that aired on television in support of proposition F was this one, which was observed for a total of 16 minutes between October 16th to 25th. The ad, which features a parody of the Eagles’ song “Hotel California,” was pulled from Youtube and the ShareBetterSF campaign website because of claims of copyright infringement. Dale Carlson, a spokesman for the campaign who contacted the Archive, wrote “We believe the ad is parody and did not constitute a copyright violation. But it had already run its course and we weren’t going to spend money on legal bills to defend an ad that was already off the air.”

In all, the Archive identified 14 unique ads opposing Proposition F that aired on TV. In the final days of the campaign, the opponents devoted airtime to this ad that calls the proposal “too extreme,” quotes from the San Francisco Chronicle, and cites high profile opponents such as Lt. Gov. Gavin Newsom, Mayor Lee. This 30-second ad aired 423 times on 10 channels in San Francisco (CNBC, CNN, FOXNEWS, KGO, KNTV, KOFY, KPIX, KRON, KTVU, MSNBC).

This review updates an earlier one issued last week focused exclusively on Airbnb ads, broadening the analysis to include all political TV ads aired from August 25th through November 3. The Archive identified ads through a number of sources, including SFGov’s Summary of Third Party Expenditures Regarding San Francisco Candidates hosted by the City of San Francisco. An audio fingerprint was created for each ad and used to find matches in some 35,000 hours of archived local station programming and cable news network shows available in the San Francisco region. The Internet Archive’s television news research library presents public opportunities to search, compare and contrast news programs in its archive. Entertainment programming is only available for select algorithmic study within its server environment.

The Internet Archive’s review of political TV ads relating to Proposition F is part of experimentation in preparation for our new Knight Foundation funded project to track political TV ads in key primary states. Stay tuned for news about our December launch.

Research by Trevor von Stein

Pro-Airbnb political TV ads air at rate of 100:1 as San Franciscans head to polls

Posted on October 29, 2015 by Nancy Watzman

For every one minute of political ads aired in favor of a contentious ballot initiative intended to further regulate Airbnb’s growing presence in the city where it is headquartered, more than 100 minutes of ads urging them to vote “no,” have aired on local San Francisco area TV stations, according to an assessment of the Internet Archive’s television archive.

Audio fingerprinting of YouTube-hosted advertising was used to identify the same ads in local station programming and cable news networks available in the region, from August 25^th through October 26^th. Sample ads can be viewed here, and metadata about their occurrences can be downloaded from this page.

Proposition F, which is backed by a coalition of unions, land owners, housing advocates, and neighborhood groups, would restrict private rentals to 75 nights per year as well as enact rules that would ensure that hotel taxes are paid and city code followed. It would also allow private party lawsuits by neighbors against private renters suspected of violating the law.

The Internet Archive found just one TV ad favoring the initiative, also appeared on the Proposition F campaign website. The Archive discovered 32 instances of this ad airing on local TV stations, for a total of 16 minutes of airplay. However, the ad, which features a parody of the song “Hotel California,” by the Eagles, (the lyrics were replaced with “Hotel San Francisco,”) was recently removed from the official website because of a claim of copyright infringement.

In contrast, in our sample range, Airbnb supporters aired more than 26 hours of ads against the initiative. One example ad, which is below, claims that the initiative would “encourage neighbors to spy on each other,” and “create thousands of new lawsuits.” This ad played at least 358 times in recent weeks, for a total of 179 minutes of airtime.

Over all, according to reports filed with the San Francisco Ethics Commission, opponents of Proposition F have reported spending $6.5 million compared to $256,000 from organizations supporting the initiative.

Of course the ad campaigns are not just limited to television. Airbnb apologized last week after it caught flack for a series of controversial bus stations and billboard ads that critics called “passive aggressive” and “whiny,” for complaining about how public institutions, such as libraries, spent their tax revenue-derived budgets.

But TV remains a key way that political operators try to influence voters. As Nate Ballard, a Democratic strategist recently said on a local newscast: “That’s how you win campaigns in California, on TV.”

research by Trevor von Stein

Get your Dem debate visualizations here

Posted on October 16, 2015 by Nancy Watzman

Hot off the internet presses, here is media analyst’s Kalev Leetaru’s visualization tool, fueled by Internet Archive data, which enables users to trace particular phrases used in broadcast news coverage in the first 24 hours after would-be presidential nominees appeared in the first Democratic debate of the 2016 election.

Scroll down and what sticks out immediately are the two subjects that captured most of the news broadcasters’ attention: “Bernie Sanders’ “damn emails” quote and guns.

When the subject came up of the controversy over Clinton’s decision to do public work from a private email server, rather than attack Clinton, Sanders defended her:

“Let me say — let me say something that may not be great politics. But I think the secretary is right, and that is that the American people are sick and tired of hearing about your damn e-mails.”

According to Internet Archive data, that sound bite aired 496 times across stations.

The other issue that grabbed attention was gun violence: Sanders, who hails from gun-friendly rural Vermont, was called to task for his vote to make it tougher to hold gun manufacturers liable when the guns they make are used in a crime. Answering a question by CNN moderator Anderson Cooper, on whether Sanders is tough enough on guns, Clinton said:

“No, not at all. I think that we have to look at the fact that we lose 90 people a day from gun violence. This has gone on too long and it’s time the entire country stood up against the NRA. The majority of our country…(APPLAUSE)… supports background checks, and even the majority of gun owners do.”

This clip aired 260 times across stations.

However, these are just the top take-aways from this massive data crunching tool. It provides a search mechanism for the user to do deeper dives into the data and discover trends across and within certain types of news broadcasts.

Leetaru’s own analysis is here, on the Washington Post’s Monkey Cage. Among his observations:

There was also variation in how much attention each network paid to each candidate (you can see for yourself using the interactive visualization). Telemundo favored Sanders with 41 percent, followed by O’Malley with 24 percent and Clinton at just 21 percent, though admittedly, they broadcast a relatively small number of excerpts. FOX Business also favored Sanders 50 percent to Clinton’s 38 percent, as did CSPAN with Sanders at 52 percent to Clinton’s 44 percent. All other networks favored Clinton, though sometimes by a relatively close margin — like CNBC (50 percent Clinton to 43 percent Sanders) or PBS affiliates (41 percent Clinton to 38 percent Sanders).

This tool is also part of the Internet Archive’s testing of technology that we’ll use in our new Knight Foundation funded project to track political TV ads in key primary states, which will launch in early December.

Dig in and have fun.

As Democratic candidates debate, Internet Archive will be gathering data

Posted on October 13, 2015 by Nancy Watzman

When Hillary Clinton and Bernie Sanders take the podium tonight along with other contenders for the Democratic presidential nomination in 2016, their debate will be televised. The Television Archive will be tracking the news coverage surrounding the debate, viewable and searchable, here.

And this tool, developed by political scientist Kalev Leetaru and fueled by Internet Archive data, allows users to see how many times a particular candidate’s name is mentioned in news coverage. Going into the debate, Hillary Clinton is getting more than twice as mentions as Sen. Bernie Sanders.

We take for granted that candidates will debate on screen, but it wasn’t always so. The faceoff between Republican Vice President Richard Nixon and Democrat U.S. Senator Jack Kennedy in 1960, 55 years ago last month, marked the first time that Americans were able to watch candidates for the nation’s highest office from the comfort of their living rooms. You can see part one of the debate here, preserved on the Archive’s servers:

The received wisdom about this famous debate was that, from this point on, candidates had to think not just about what they said on the campaign stump, but how they looked. This could make a huge difference in how the public and the media perceived who “won” the debate. Nixon looked tired and like he needed a shave. Kennedy looked healthy and vibrant. Those who listened on the radio thought Nixon won.

“It’s one of those unusual points in the timeline of history where you say things changed very dramatically–in this case, in a single night,” Alan Schroeder, a media historian and associate professor at Northeastern University, told Time Magazine in 2010.

Here’s part II of the Kennedy-Nixon 1960 debate:

We don’t know yet who the perceived winner of tonight’s debate will be. The Internet Archive’s data will provide one way to evaluate this. Stay tuned.

Internet Archive Blogs

A blog from the team at archive.org