A Microsoft Director of Applied Science Brent Hecht privately dismissed AI training data scraping as the largest theft of labor in human history,
according to newly unsealed Manhattan Federal Court filings made public on September 17, 2026, in a landmark copyright infringement lawsuit brought by major news organizations.
Court Filings Reveal Internal Microsoft Admissions on AI Scraping
The landmark copyright infringement lawsuit playing out in Manhattan Federal Court between news outlets and tech giants Microsoft and OpenAI has brought to light internal executive warnings that undercut the defendants’ public legal defense. Brent Hecht, Microsoft’s Director of Applied Science, wrote in internal company documents shortly after the litigation commenced that the industry-wide practice of scraping the open web to train large language models would soon be viewed by millions as an astonishing theft of unprecedented proportions.
In a January 2023 internal memo, he characterized the automated harvesting of uncompensated human labor as the largest theft of labor in human history,
adding that almost no one intended for their creative work to be utilized in such a fashion or expected to receive compensation for it, according to reporting on the court disclosures.
“[M]illions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions.”
Brent Hecht, Director of Applied Science at Microsoft
Scale of Content Acquisition and Bypassed Paywalls
The unsealed documents reveal precise figures concerning the volume of material ingested to train the underlying models. OpenAI mid-training datasets contain more than 91,692 separate copies of articles sourced from The New York Times, the New York Daily News, and the Center for Investigative Reporting. A separate training set constructed from Common Crawl, a public nonprofit web archive utilized widely across the artificial intelligence sector, held upward of 2 million documents pulled directly from nytimes.com.
Legal representatives for the plaintiff newspapers assert that the tech companies engaged in illicit copying of billions of web pages not merely for direct model training, but for internal horse trading
deals where datasets were traded among technology firms rather than licensed from copyright holders. Specific figures unredacted from the filings point to more than 3.9 million copies acquired from The New York Times and 7.3 million copies taken from the Orange County Register, the New York Daily News, The Chicago Tribune, the Orlando Sentinel, the Denver Post, and the Mercury News.
The filings also outline technical methods used to bypass restricted content. In a chat excerpt cited by plaintiffs’ lawyers, OpenAI researcher Nick Ryder informed CEO Greg Brockman about a technique for circumventing The New York Times paywall. Brockman’s documented response consisted of two words: ah nice.
Internal Warnings Over Referral Traffic Cannibalization
Beyond initial training data acquisition, Microsoft’s internal communications tracked the economic fallout the company’s own products inflicted on news publishers. In a January 2024 presentation, Hecht warned that the Microsoft Copilot answer engine directly cannibalized the referral traffic it was designed to supplement. Internal data demonstrated that click-through rates to New York Times web pages plummeted by as much as 93% when compared against traditional Bing search result patterns.

Hecht labeled this dynamic a doom loop,
noting that it threatened the performance of the tech companies’ own models alongside the economic viability of the open web.
Legal Repercussions and Defense Posture
The unsealed record strikes at the core of the fair use doctrine advanced by Microsoft and OpenAI. Legal counsel for the publishers argues that internal admissions acknowledging theft dismantle any claim of good-faith operation, while evidence of traffic reduction undermines arguments concerning market harm. Steven Lieberman, an attorney representing the Daily News in the litigation, noted the significance of the unsealed disclosures, stating that the public can finally review how the companies regarded the fairness of their behavior.
Microsoft has distanced itself from the statements. A company spokeswoman asserted that Hecht’s remarks represented an individual employee’s perspective rather than official legal analysis or corporate policy, maintaining that Copilot use aligns with copyright protections.
The litigation, which consolidates actions from news publishers, The Authors Guild, and best-selling writers following an initial complaint filed in December 2023 and a federal judge’s ruling in April 2025, moves forward in Manhattan Federal Court as plaintiffs pursue summary judgment alongside separate sanctions against OpenAI for alleged evidence destruction.
Keep reading