Microsoft Executive Called AI Scraping "Theft" in Unsealed Court Filings

Unsealed court filings show Microsoft and OpenAI executives privately called AI training on news content "theft" and an "existential threat" to publishers.

maisiekooc
Maisie Morrison

AgentLocker Editor

AI News
Microsoft Executive Called AI Scraping "Theft" in Unsealed Court Filings

Court filings unsealed this week add new detail to a three-year-old lawsuit filed by The New York Times against OpenAI and Microsoft. The case centers on whether the companies broke copyright law by training AI models on the paper's articles without permission.

The newly unredacted material includes internal messages from executives at both companies. Some of those messages describe the scraping of news content in blunt terms.

Microsoft's director of Applied Science, Brent Hecht, wrote in a January 2023 memo that the practice amounted to "an astonishing theft of unprecedented proportions." He called it "the largest theft of labor in human history."

What the Filings Allege

The filings say OpenAI's mid-training datasets contain more than 91,692 copies of articles from The New York Times, Daily News, and Center for Investigative Reporting. A separate dataset built from Common Crawl included over 2 million documents from nytimes.com alone.

The documents also describe two internal data-sharing efforts between the companies, called Project Taxi and Project Mango. Filings say the Project Mango dataset alone contains at least 160,903 unique works from the news publishers involved in the case.

Messages cited in the filing suggest OpenAI staff discussed a method to get around the Times' paywall. When a researcher told OpenAI President Greg Brockman about the workaround, Brockman reportedly replied, "ah nice."

The filings also claim researchers worked to strip copyright notices from training data before it reached the models. The stated reason was to avoid the models repeating those notices to users.

Business Risk Flagged Internally

Separate from the copying claims, the filings include Microsoft's own performance data. That data reportedly shows Copilot's presence in Bing search results caused click-through rates to The New York Times site to fall by as much as 93% compared with traditional search results.

Hecht described this trend as a "doom loop" in an internal presentation from January 2024. He warned it could hurt both Microsoft's models and the wider web over time.

Microsoft CEO Satya Nadella addressed the topic in a deposition earlier this year. He said paywalled content "should be licensed by anyone who wants to use it" for training or grounding AI systems.

Nadella added that if he had known OpenAI trained on paywalled material without permission, he would have pushed OpenAI to retrain its models.

OpenAI's head of ChatGPT, Nick Turley, wrote separately that AI products pose an "existential threat" to publishers. He described chatbots as "largely substitutive" for news content and said that substitution would grow as the models improve.

Much of the newly unsealed material comes from the Times' own legal brief rather than the underlying exhibits, which remain sealed. That means some quotes are presented without their full original context.

The case remains ongoing. Earlier this month, the Trump administration filed a brief supporting OpenAI's position that training on copyrighted material qualifies as fair use under existing law.

From our research desk
AI Jobs Automation Index
Which jobs are AI tools targeting most? We mapped 3,400+ AI tools to real job functions — with BLS employment & salary data.
Explore the index
maisiekooc

Written by

Maisie is a news writer at Agent Locker, covering the latest developments in artificial intelligence, emerging technology and the companies shaping the future.

Discover AI Agents