Unsealed court filings reveal Microsoft internal criticism of AI data scraping
Newly unsealed court documents show that a Microsoft executive privately described OpenAI's data scraping practices as the 'largest theft of labor in human history.' The filings also reveal internal concerns at both companies regarding the potential impact of AI on the news industry.
First reported 1 hour ago · latest update 1 hour from now
Newly unredacted court filings have revealed internal communications from Microsoft and OpenAI regarding the development of generative artificial intelligence. In documents submitted as part of an ongoing copyright lawsuit, a senior Microsoft executive privately characterized the practice of scraping data to train AI models as the largest theft of labor in human history.
The filings indicate that leadership at OpenAI also expressed concerns regarding the impact of their technology on the media industry. Internal discussions suggested that the AI models being developed posed an existential threat to the publishers and journalists whose work was utilized to train the systems.
The legal documentation outlines allegations that the companies engaged in systematic methods to acquire training data. These claims include assertions that the firms bypassed paywalls to access content, utilized mass scraping techniques to build datasets, and deliberately removed copyright notices from the materials used in the training process.
Much of the information regarding these internal sentiments originates from briefs filed by the plaintiff in the case. While these filings provide insight into the private perspectives of company leadership, the underlying exhibits containing the full context of these communications remain sealed by the court.
The lawsuit, which has been active for three years, centers on the question of whether the use of copyrighted material for AI training constitutes a violation of intellectual property law. While the legal status of these practices remains unsettled, courts have frequently leaned toward the argument that training AI models qualifies as fair use. This legal doctrine allows for the use of copyrighted material under specific circumstances, and its application to generative AI remains a central point of contention in the ongoing litigation.
Citations · 3 reports from 3 outlets
Tap a citation to read it above, right here on T.A.M.
Microsoft exec called AI the ‘largest theft of labor’ in history, court records show - washingtonpost.com
Microsoft exec called AI the ‘largest theft of labor’ in history, court records show washingtonpost.com
1 hour from nowMicrosoft exec called AI scraping the “largest theft of labor in human history”
Microsoft, OpenAI emails reveal fear of AI “doom loop” killing news orgs.
1 hour ago · Ashley BelangerMicrosoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal
Newly unsealed court filings show Microsoft privately called OpenAI's data practices "theft" while both companies scraped paywalled Times content, built datasets from it, and warned internally it would gut publishers.
1 hour ago · Rebecca Bellan