Internal Communications Reveal Stark Views on AI Data Practices
The ongoing lawsuit by The New York Times against OpenAI and Microsoft has unearthed internal communications that paint a stark picture of the companies' views on the data used to train AI models. Revelations from legal filings indicate that senior figures within both Microsoft and OpenAI have expressed grave concerns, even alarm, regarding the methods used to acquire training data, particularly concerning the scraping of copyrighted material. A Microsoft director, identified in court documents, described the practice of AI scraping as ".the largest theft of labor in human history." This strong condemnation, coming from within a company that is a major partner and investor in OpenAI, suggests a deep internal conflict or at least a significant degree of public relations management surrounding the contentious issue of AI training data. The statement implies that the vast amounts of human-created content — from articles and books to code and art — that fuel large language models (LLMs) are being exploited without adequate compensation or permission. The director's phrasing, ".the largest theft of labor in human history,", elevates the debate beyond mere copyright infringement to a systemic exploitation of human creative and intellectual output on an unprecedented scale. Adding to this critical perspective, the head of OpenAI, Sam Altman, is quoted in the same legal filings as branding ChatGPT an ".existential threat" to publishers. This characterization is particularly telling, as it acknowledges the disruptive potential of generative AI on industries that rely on original content creation and distribution. For publishers, ChatGPT and similar models can generate content that directly competes with their offerings, potentially siphoning off readership and advertising revenue. The threat is existential because it questions the fundamental business models upon which news organizations and other content creators have operated for decades. The revelations stem from legal briefs submitted as part of The New York Times' lawsuit, which accuses OpenAI and Microsoft of using millions of its copyrighted articles without permission to train AI models. The lawsuit seeks to hold the companies accountable for allegedly infringing on the newspaper's intellectual property rights and seeks damages, though the exact amount is not specified. The Times argues that its content was scraped from its website and used to build AI systems that can now produce outputs that directly mimic or substitute for the original reporting, thereby undermining the value of its journalistic work.
