NYT Alleges OpenAI Concealed Key Evidence
The New York Times has escalated its copyright infringement lawsuit against OpenAI, filing a motion for sanctions and accusing the artificial intelligence lab of deliberately hiding crucial evidence. The publisher asserts that OpenAI withheld tools and datasets that could have identified instances where ChatGPT generated content derived from copyrighted journalistic material. This development marks a significant turning point in the ongoing legal battle, suggesting a deliberate attempt by OpenAI to obscure the extent of its alleged infringement. The core of the lawsuit revolves around whether OpenAI's large language models (LLMs), particularly ChatGPT, were trained on vast amounts of copyrighted material without permission. The New York Times, along with other publishers, claims that OpenAI's models reproduce protected text, effectively using their journalism to train a product that competes with them. The motion for sanctions, filed in the Southern District of New York, seeks to penalize OpenAI for what the Times describes as "extraordinary misconduct." According to the Times' filing, OpenAI failed to produce specific evidence related to the datasets and tools used for training, especially those that could pinpoint copyrighted journalistic content within ChatGPT's outputs. This alleged concealment is not merely a procedural hiccup; the Times argues it directly obstructs their ability to prove the central claims of their case. Without access to these specific analytical tools and data, it becomes significantly more challenging for the plaintiffs to demonstrate the direct lineage of copyrighted text appearing in AI-generated responses.
