News

Unsealed Filings Show Microsoft Labeled OpenAI's Data Practices as 'Theft' of Publisher Content

A batch of unredacted court filings has surfaced, offering a rare glimpse into how Microsoft internally assessed OpenAI's approach to acquiring training data. According to the documents, a Microsoft executive described the mass scraping of copyrighted material as "the largest theft of labor in human history"—a stark assessment made while both companies were actively using paywalled New York Times content to build AI datasets.

The filings indicate that Microsoft officials recognized the potential damage such practices could inflict on publishers. Internal warnings reportedly stated that the approach could "gut publishers" by undermining the economic model that supports journalism. The documents emerged as part of ongoing litigation related to AI training data practices.

The unsealed materials highlight the tension between AI developers and content creators over how large language models are trained. While AI companies have argued that training on publicly available data falls under fair use, publishers and creators contend that scraping paywalled material without compensation constitutes copyright infringement.

The filings shed light on a period when major technology companies were racing to amass datasets for AI development, often without clear legal frameworks or compensation structures in place. The internal Microsoft communications stand in contrast to the company's public-facing positions on AI and intellectual property.

Sources