Home
Several US Media Outlets Block Internet Archive's Wayback Machine Crawler to Prevent AI Training Abuse
As reported by Wired, several major media outlets and platforms—including The New York Times, Reddit, and USA Today's parent company—have recently blocked the Internet Archive's Wayback Machine. The move is intended to prevent AI companies from using the archive tool to indirectly scrape copyrighted material for model training.

The irony of blocking while benefiting
Ironically, USA Today's recent in-depth investigation into immigration policy statistics relied on historical data preserved by the Wayback Machine. Yet the media group's spokesperson stated they have fully blocked all crawlers, including ia_archiverbot, to address the growing risk of AI infringement.
How media organizations are applying restrictions
At least 23 mainstream news websites have now implemented restrictions:
Complete blocking: The New York Times and Reddit have directly blocked the Wayback Machine's dedicated crawler.
Interface filtering: The Guardian has not fully blocked crawlers but has excluded its content from the Internet Archive's API and filtered its search interface, making it extremely difficult for users to access historical archives.
In response to these blocking actions, over 100 active journalists, including Rachel Maddow, have co-signed a letter of support to the Electronic Frontier Foundation (EFF). They describe the Wayback Machine as an "indispensable tool" for fact-checking, tracking changes in the behavior of powerful institutions, and preserving digital history.
Publishers argue that AI companies training on the Internet Archive's vast data violates copyright law and directly competes with them. However, Internet Archive director Mark Graham warns that the ongoing closure of public web content is seriously undermining society's ability to understand historical truths and conduct public oversight. If this trend continues, a large portion of early digital records may be at risk of complete loss.
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500
As reported by Wired, several major media outlets and platforms—including The New York Times, Reddit, and USA Today's parent company—have recently blocked the Internet Archive's Wayback Machine. The move is intended to prevent AI companies from using the archive tool to indirectly scrape copyrighted material for model training.

The irony of blocking while benefiting
Ironically, USA Today's recent in-depth investigation into immigration policy statistics relied on historical data preserved by the Wayback Machine. Yet the media group's spokesperson stated they have fully blocked all crawlers, including ia_archiverbot, to address the growing risk of AI infringement.
How media organizations are applying restrictions
At least 23 mainstream news websites have now implemented restrictions:
Complete blocking: The New York Times and Reddit have directly blocked the Wayback Machine's dedicated crawler.
Interface filtering: The Guardian has not fully blocked crawlers but has excluded its content from the Internet Archive's API and filtered its search interface, making it extremely difficult for users to access historical archives.
In response to these blocking actions, over 100 active journalists, including Rachel Maddow, have co-signed a letter of support to the Electronic Frontier Foundation (EFF). They describe the Wayback Machine as an "indispensable tool" for fact-checking, tracking changes in the behavior of powerful institutions, and preserving digital history.
Publishers argue that AI companies training on the Internet Archive's vast data violates copyright law and directly competes with them. However, Internet Archive director Mark Graham warns that the ongoing closure of public web content is seriously undermining society's ability to understand historical truths and conduct public oversight. If this trend continues, a large portion of early digital records may be at risk of complete loss.
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage











