Technology & Science

Unsealed Filings Show Microsoft & OpenAI Flagged Their Own AI Data Harvest as "Largest Labor Theft"

On 19 September 2026 the New York Times unsealed discovery papers revealing that, as early as January 2023, senior Microsoft and OpenAI staff internally branded their web-scraping AI training strategy a self-described "doom loop" and acknowledged it would strip publishers’ traffic while copying millions of pay-walled articles.

By Underlines Team

Focusing Facts

  1. Microsoft memo quoted Director of Applied Science Brent Hecht calling the mass copying an “astonishing theft of unprecedented proportions … the largest theft of labor in human history.”
  2. Internal telemetry showed Bing Chat click-throughs to NYT pages were 87-93 percent lower than from standard Bing Search.
  3. Through the Microsoft-run "Project Mango" crawler, at least 160,903 distinct works from the plaintiff publishers were copied for OpenAI training between 2019–2022.

Context

The tension echoes Google’s 2004–2015 Google Books litigation—another moment when a tech firm argued transformative fair use after bulk scanning millions of texts—but now the scale is orders larger and the substitution effect immediate. Since the 1710 Statute of Anne, copyright has expanded whenever new reproduction technology (the rotary press in 1843, the phonograph in 1877, Napster in 1999) threatened creators’ revenue model; these filings suggest LLMs are the next inflection. They expose a systemic pattern: platforms first ingest free cultural labor to bootstrap value, then hollow out the originators’ market—what economists call enclosure of digital commons. Whether courts bless or curb this practice will shape knowledge-work incentives for decades; if unchecked, AI developers might deplete the very data ecosystems their models need, repeating a 100-year boom-and-bust cycle of content commodification and legal catch-up.

Perspectives

Tech watchdog and critical tech press

e.g., RocketNews, Thurrott.comThe unsealed filings prove Microsoft and OpenAI deliberately committed large-scale "theft" of news content, creating an AI-driven “doom loop” that undercuts publishers and violates copyright. These outlets lean on incendiary language and worst-case forecasts to attract attention and vilify big-tech actors, glossing over the unresolved legal questions around fair use that courts will ultimately decide.

Corporate defense from Microsoft & OpenAI

statements highlighted in mainstream and financial press such as The Indian Express, MoneyControlCompany spokespeople argue that scraping publicly available material is a transformative fair use comparable to search indexing, insisting their AI tools don’t substitute for original journalism and therefore do not infringe copyright. These carefully crafted statements are motivated by protecting multi-billion-dollar AI investments and minimizing legal exposure, so they dismiss incriminating employee emails and emphasize legal technicalities over publishers’ economic harms.

Like what you're reading?

Create a free account to read 5 articles every week. No credit card required.

Share

Related Stories