SDSignal Desk

Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal

Sep 17, 2026, 12:46 PM · TechCrunch

Image: TechCrunch

TechCrunch’s unredacted-filing read adds scale: 91,000+ mid-training copies of plaintiff works, 2M+ nytimes.com Common Crawl docs, Project Mango’s 160,903 unique publisher works, and Brockman’s “gazillions” motivation line.

Why it matters

Same lawsuit week, sharper inventory. The Times brief—quoting still-sealed exhibits—puts numbers on the scrape.

Hecht’s January 2023 memo: largest theft of labor in human history. Copilot answer-engine CTR for NYT down as much as 93% versus Bing search. Nadella says he’d have forced a retrain if he’d known OpenAI trained on paywalled material. Mid-training sets alone hold more than 91,692 copies of Times, Daily News, and CIR works.

Trump administration briefing for unlicensed training sits awkwardly next to defendants’ own substitution talk.

From the desk

We’re stacking this TechCrunch pass with the Ars account because the counts change the texture of the fight.

Project Taxi and Project Mango data swaps, WebText’s news-heavy tilt, copyright-notice stripping so models wouldn’t spit notices back—these are process choices, not accidents. Ryder-to-Brockman paywall hack banter remains the character exhibit. Hecht’s doom-loop deck warned the product threatens its own suppliers; Turley called chatbots largely substitutive, period.

Caveat the story itself flags: much arrives via the Times brief, not public exhibits. Still, the quoted lines cut against market-harm prongs of fair use after months of judicial sympathy for trainers—and after the administration’s brief backing OpenAI.

Useful models trained on a dead press aren’t a win. I’m watching whether these figures force settlement math before a jury ever sees the cartoons Microsoft put in its own doom-loop slides.

Context

TechCrunch, September 17, 2026. Companion reporting on the same unsealed NYT v. OpenAI/Microsoft materials covered by Ars.

Who feels it

Litigators and judges
Quantitative scrape inventories and CTR deltas become central fair-use exhibits.
Newsrooms
Stronger public proof that AI firms modeled publisher harm internally while shipping substitutive products.
Investors
Licensing or adverse judgments could reprice training-data risk across the sector.

What to watch

  1. Unsealing of underlying exhibits beyond the brief’s quotations.
  2. OpenAI/Microsoft formal responses if they break silence.
  3. Parallel suits citing these admissions.

Read the original

Continue at the source.

TechCrunch

Companies: OpenAI, Microsoft