The AI Scraping War has been begin.
For two years, AI companies scraped the open web largely unchallenged, building the training data behind chatbots and search tools out of journalism, reference works, and forums without asking permission or paying for it. That era appears to be ending — not through a single dramatic event, but through a rapid accumulation of lawsuits, blocked crawlers, ultimatums, and collapsing referral traffic that’s forced publishers into open confrontation with the AI industry they’ve spent two years trying to hold at bay.
The scale of what’s now at stake is genuinely large: journalism jobs are being cut in the thousands, one of the internet’s largest infrastructure companies has threatened to cut Google out of search results entirely, and a coalition of major news organizations has formally accused a nonprofit web archive of enabling mass copyright infringement. Here’s the full shape of the fight as it stands.
The Lawsuits Piling Up
Legal action against AI companies has escalated sharply over the past year. In late May 2026, publishers that collectively own and operate nearly 400 newspapers sued OpenAI and Microsoft, alleging their content was scraped to build products like ChatGPT and Microsoft Copilot without permission or compensation. The complaint argues these products have generated billions of dollars in market value for the two companies, while “not a cent of it has gone” to the publishers whose work made those products possible.
That case joined an already crowded legal field. The New York Times and Chicago Tribune have sued Perplexity AI as part of the same broader copyright conflict. Reddit has sued Perplexity and other companies over alleged data scraping. Britannica and Merriam-Webster separately sued Perplexity over scraping of their reference content. Each of these cases is being fought on similar legal ground — whether training AI systems on copyrighted material without a license constitutes infringement — but the sheer number of plaintiffs, spanning newspapers, reference publishers, and social platforms, signals just how widely this conflict has spread across different types of content owners.
Publishers vs. Common Crawl
One of the sharpest recent escalations moved beyond individual AI companies entirely, targeting the data infrastructure underneath the industry. On June 3, 2026, the trade group Digital Content Next sent a cease-and-desist letter to Common Crawl — a web-archiving nonprofit whose datasets are widely used to train AI models — on behalf of the Associated Press, the New York Times, NBC Universal, Bloomberg, NPR, and Fox.
The letter makes two specific demands: that Common Crawl stop scraping protected content going forward, and that it remove member content already sitting inside datasets that AI labs use for training, including paywalled and subscriber-only articles. It accuses Common Crawl of infringing copyrighted content by creating and distributing these datasets while knowing AI companies would use them to reproduce protected material. Targeting Common Crawl directly, rather than just the AI companies that use its datasets, represents a meaningfully different legal strategy — going after a shared upstream resource rather than fighting each AI company’s use of scraped content individually.
The Traffic Collapse Behind the Fight
The financial motivation behind this legal escalation shows up clearly in traffic data. Google search referrals to publishers fell roughly a third globally in the year leading up to November 2025, and a study from Ahrefs measured a 58% click-through-rate drop on top-ranking pages where Google’s AI Overviews feature appears above the traditional search results. In practice, that means AI systems are increasingly answering user questions directly using scraped publisher content, without sending the reader to the original source at all.
The human cost of that shift is already visible in employment data. More than 3,400 journalism jobs were cut across the US and UK in 2025 alone, and 2026 is reportedly running ahead of that pace, with deep newsroom cuts already reported at the Washington Post and Nexstar.
Cloudflare’s Ultimatum to Google
Perhaps the most consequential recent move came from infrastructure rather than publishing directly. Cloudflare, which sits in front of roughly one-fifth of all websites globally, gave Google an ultimatum: change how it handles AI scraping of publisher content indexed through Google Search, or face being cut off from indexing the publishers Cloudflare protects, beginning in September.
Cloudflare’s leverage here comes from scale and technical positioning — as the content delivery network sitting between huge portions of the web and its visitors, it’s uniquely positioned to enforce this kind of ultimatum at an infrastructure level rather than site by site. The company has separately disclosed blocking 416 billion AI bot requests since July 2025, and in the same month shifted its default policy to “permission-by-default” for AI crawlers, alongside launching a “pay-per-crawl” licensing model with the Associated Press, Time, The Atlantic, and Reddit signing on as launch partners.
Scraping Has Become Its Own Industry
One of the more striking developments is how thoroughly AI scraping has evolved from an unwanted side effect of the open web into what’s now being described as its own standalone media business. Bots ingest published reporting, AI models repackage that content inside chatbots and search overviews, and the original publisher frequently never sees the reader who consumed their work at all — a “scrape-summarize-monetize” loop that industry observers now describe as the dominant news distribution channel for a large share of readers.
Compounding the problem for publishers, Matt Rogerson, the Financial Times’ director of global public policy and platform strategy, has noted that even after two years of publishers trying to close down every loophole in their website security, real gaps remain — including scraping-for-hire platforms that specialize in getting behind paywalls specifically to extract protected content on behalf of paying clients.
Why the Internet Archive Became a Liability
Even institutions built around open information access have been caught in the crossfire. The Internet Archive’s Wayback Machine, which preserves snapshots of webpages as part of its mission to keep the web archived and accessible, has become an unexpected vulnerability for publishers — because AI bots scavenging the web for training data can use archived snapshots as a backdoor around a publisher’s own paywalls and access controls.
The Guardian’s head of business affairs and licensing has said that when the outlet reviewed its own access logs to see who was extracting its content, the Internet Archive itself showed up as a frequent crawler. In response, outlets including The Guardian and the New York Times have begun scrutinizing their relationship with digital archives, treating them as a potential AI-scraping backdoor rather than a purely benign preservation tool.
Where This Might Actually Be Heading
Despite the intensity of the conflict, there are genuine signs of a shift toward negotiated licensing rather than pure confrontation. The Financial Times was the first UK-based publisher to strike a licensing deal with OpenAI back in 2024, and Rogerson believes 2026 will bring a broader reset as major AI companies adjust their approach specifically to reduce future legal exposure. He points to a growing number of institutions and corporations taking AI summarization licenses, driven by their own recognition that AI systems are only genuinely valuable to their businesses if the underlying content is accurate and comes from sources they can trust.
Standardization efforts are also gaining real traction. Work that began in 2025 on frameworks like the IAB Tech Lab’s CoMP framework and RSL’s licensing standards is expected to mature further in 2026, alongside the IAB’s own proposed “AI Accountability for Publishers Act” legislation, introduced in February 2026 specifically to address large-scale scraping of publisher content. Industry observers note a level of collective publisher action behind these efforts that stands in sharp contrast to the earlier Facebook-Google digital advertising duopoly years, when publishers largely failed to organize a unified response. That said, not every publisher is waiting on collective solutions — larger outlets like News Corp and the New York Times remain big enough to negotiate individual licensing terms on their own.
Conclusion
What began as a diffuse, largely uncontested scraping of the open web has become an active, multi-front war — lawsuits against the largest AI companies, a direct legal challenge to the data infrastructure underneath the entire industry, and one of the internet’s most important infrastructure providers threatening to cut off Google itself. Whether 2026 actually becomes the “reset” some publishers are hoping for, with standardized licensing replacing unauthorized scraping, will likely depend on whether AI companies conclude that paying for trustworthy content is cheaper than continuing to fight lawsuits from nearly every major publisher in the industry at once.