Microsoft Exec Labels OpenAI’s AI Data Scraping ‘Largest Theft of Labor’

3 min readSources: TechCrunch

A Microsoft exec called OpenAI’s scraping of paywalled news ‘the largest theft of labor in human history.’

Why it matters: This highlights major legal and ethical challenges around AI training data and publisher rights. Legal professionals must track evolving IP claims and AI content usage disputes.

  • On September 4, 2026, Seattle Times and Newsday sued OpenAI and Microsoft for unauthorized scraping of paywalled articles to train AI models like ChatGPT.
  • A Microsoft executive described OpenAI’s scraping of paywalled journalism as 'the largest theft of labor in human history' in internal documents made public amid litigation.
  • Publishers argue AI models can reproduce or closely paraphrase their work while omitting copyright notices and bylines.
  • A coalition representing nearly 400 newspapers filed a related lawsuit in June 2026; The New York Times filed suit in December 2023 over similar claims.

Internal documents quoting a Microsoft executive reveal sharp criticism of OpenAI's use of paywalled journalism content for AI training, calling it 'the largest theft of labor in human history.' These documents surfaced amid lawsuits accusing OpenAI and Microsoft of scraping protected news content without authorization.

On September 4, 2026, the Seattle Times and Newsday lawsuit was filed, alleging that OpenAI and Microsoft scraped hundreds of thousands of paywalled articles to train AI tools like ChatGPT and Microsoft Copilot. The complaint claims the AI models can reproduce or closely paraphrase reporting while altering or omitting bylines, titles, and copyright information.

Publishers seek monetary damages and demand the destruction of datasets and AI models containing their copyrighted material. This lawsuit joins a series of similar suits, including one by a coalition of nearly 400 newspapers represented by the News Media Alliance, filed in June 2026, and an earlier December 2023 New York Times lawsuit alleging unauthorized use of millions of articles.

These cases expose ongoing tensions between AI developers who require vast datasets for training and publishers defending their intellectual property and revenue models. The Microsoft executive's comment and the active litigation underscore deep concerns about whether AI companies can rely on publicly accessible or paywalled content without consent.

Currently, all these lawsuits are ongoing. Legal rulings on these issues will likely set key precedents on the boundaries of AI training data usage and how copyright law applies to automated content scraping.

By the numbers:

  • September 4, 2026 — Seattle Times and Newsday filed their lawsuit.
  • Nearly 400 newspapers — Coalition suing OpenAI and Microsoft as of June 2026, represented by News Media Alliance.
  • December 2023 — New York Times initiated similar litigation.

Yes, but: The Microsoft executive’s statement comes from internal documents released during litigation and may reflect internal disagreement; public company statements on the issue remain more measured.

What's next: Litigation is ongoing; courts are expected to issue rulings on these AI training data copyright claims in late 2026 and early 2027, potentially reshaping AI content use policies.