Over 340 Local Publishers Block Internet Archive: the Wayback Machine Loses Access
Internet Archive
Internet Archive and its Wayback Machine are seeing increasingly restricted access to the content of news publishers: according to an update from a Nieman Lab investigation, more than 340 local U.S. outlets now limit the ability of the non-profit archive to store their articles. This is not a sudden development: in January, Nieman Lab documented how major publishing groups, including the New York Times, The Guardian, and USA Today Co., began to block the archive's crawlers.
We have previously discussed it, and in the five months since, the phenomenon has expanded.
The updated sample counts 382 outlets that block at least one bot related to Internet Archive, of which 342 are local; 93% are based in the United States, with the rest distributed across ten countries. In January, the number of sites was 241, about 80% associated with USA Today Co. (formerly Gannett); since then, an additional 141 have been added. Many of these outlets belong to five of the seven largest local publishers in the country: in addition to USA Today Co., they include McClatchy, Advance Local, MediaNews Group, and Tribune Publishing. The latter two are both controlled by Alden Global Capital, the hedge fund known for acquiring newspapers and reducing their resources.
Reasons for the Block
The recurring reason is the fear that artificial intelligence companies will use the Wayback Machine as a shortcut to collect content to train their models. It is worth noting that this concern remains hypothetical: no publisher has confirmed to Nieman Lab that an AI company has actually pulled its articles from the archive.
Positions, however, are not homogeneous. The Atlantic has adopted what its spokesperson Anna Bross describes as an "aggressive" blocking policy, with exclusion as the default setting, and has been collaborating with Cloudflare since last summer; CEO Nick Thompson explained that preventing scraping helps maintain negotiating power in licensing agreements with large AI companies. In contrast, the Baltimore Banner remains open to the appearance of its articles in chatbots and allows crawlers from ChatGPT and Claude, but blocks Internet Archive for another reason. "The threat is certainly not Internet Archive," technical lead Biswajit Ganguly stated: for the Banner, the point is to ensure that AI-based products trace back to the original source instead of attributing the content to the aggregating sites.
The blocking is not limited to local publishing. Condé Nast has extended restrictions to Vogue, The New Yorker, Wired, Pitchfork, Vanity Fair, and Bon Appétit, while the Brazilian outlet Folha de S.Paulo is also among the international publications.
What is Lost
The practical effect falls on those who use these archives every day: researchers, historians, journalists, and fact-checkers, who often rely on the Wayback Machine to recover articles that have disappeared from closed or reorganized outlets. The losses are not theoretical: in 2024, thousands of articles from the Daily Hampshire Gazette and the Greenfield Recorder vanished during a CMS change, while in 2022, the archive of the weekly The Hook went offline, taking over 22,000 stories with it.
Alternatives do exist, but they are paid: publishers have long licensed their content to commercial archives such as ProQuest and LexisNexis, accessible through libraries, universities, or individual subscriptions. And this, of course, represents a possible economic incentive behind the closure towards a free service. Mark Graham, the founder of Wayback Machine, reminded that the archive's terms of use allow access to collections only for study or research purposes, and that dialogue with publishers remains open. Meanwhile, a petition is active that asks outlets to resume collaboration with the archive.
However, the controversy rests on a much "older" dispute. As noted by Meredith Broussard, a professor at New York University, it is "the same battle that everyone has been fighting with Internet Archive since its inception", between those who believe that information should be free and those with different priorities; AI, in this view, is merely "the trigger for the latest skirmish." Her conclusion on the duration of archives is quite terse: "Anyone who told you that the internet is forever has lied to you."