Music Publishers Allege Widespread IP Theft

Sony Music Publishing and Warner Chappell have filed a lawsuit against AI company Anthropic, along with its CEO Dario Amodei and co-founder Benjamin Mann. The suit, filed on August 28, centers on accusations that Anthropic engaged in a systematic and illegal campaign to acquire and use copyrighted song lyrics for training its large language models. This legal action is notable because it does not seek to establish new legal precedents but rather to apply existing rulings to a new set of alleged infringements.

The core of the complaint revolves around the alleged acquisition of millions of song lyrics through illicit means, mirroring a previous legal entanglement involving Anthropic and copyrighted books. In September of the previous year, a federal judge ruled in the 'Bartz case' that while the act of training an AI model on copyrighted text itself might be legal, the method of obtaining that training data through piracy was not. Anthropic subsequently settled that case for a reported $1.5 billion after admitting that Benjamin Mann had personally downloaded over five million books from Library Genesis. Furthermore, Anthropic staff were accused of acquiring an additional two million books from Pirate Library Mirror in 2022.

Sony and Warner's current lawsuit explicitly cites these same alleged piracy channels, now linking them to the acquisition of lyric datasets from MusixMatch and LyricFind. Instead of asking courts to determine the legality of AI training on copyrighted material, the music publishers are leveraging the established legal finding from the Bartz case. They contend that Anthropic's actions constitute copyright infringement based on the illegal procurement of these lyrics, similar to how the book downloads were treated.

An illustration depicting a legal gavel striking a stack of sheet music and books

Statutory Damages and Potential Financial Fallout

The legal strategy employed by Sony and Warner appears designed to maximize financial penalties. Copyright law allows for statutory damages ranging from $150,000 per infringed work. Given the scale of the alleged piracy, involving millions of song lyrics, the potential financial damages could significantly surpass the $1.5 billion settlement paid in the book piracy case. The exact number of songs Anthropic is accused of pirating will determine the final damages, but the music publishers are clearly aiming for a substantial financial judgment.

This lawsuit is not just about recouping financial losses; it is also a clear signal to the burgeoning AI industry about the consequences of intellectual property theft. The AI companies have increasingly come under scrutiny for their data acquisition practices, with many content creators and rights holders arguing that their works have been used without permission or compensation to build powerful AI models. The music industry, in particular, has been vocal about the unauthorized use of its vast catalog of songs.

Anthropic, founded by former OpenAI researchers, has positioned itself as a responsible player in the AI space, emphasizing safety and ethical development. However, these allegations of direct involvement in piracy, particularly by senior leadership, cast a shadow over that image. The company's previous settlement over book piracy suggests a pattern of behavior that now appears to be repeating with musical content. The lawsuit names both the company and its key executives, indicating a desire to hold individuals accountable as well as the corporate entity.

Broader Implications for AI Training Data

The outcome of this lawsuit could have far-reaching implications for how AI models are trained and the legal frameworks governing data acquisition. If Sony and Warner are successful in leveraging the precedent set in the Bartz case, it could embolden other copyright holders to pursue similar legal actions against AI companies. This would force AI developers to be far more rigorous in their data sourcing, potentially leading to increased costs for licensing or the development of entirely new, legally compliant datasets.

The question of whether AI training itself is fair use remains a complex and highly debated topic. However, this lawsuit bypasses that fundamental question by focusing on the methods of data acquisition. By proving that Anthropic obtained its training data through piracy, Sony and Warner are sidestepping the need for a court to rule on the legality of AI training on copyrighted material. This approach is a direct application of a previously established legal standard to a new set of alleged violations.

Anthropic's defense will likely hinge on disputing the scale of the piracy, the number of infringed works, or potentially arguing that the specific datasets acquired were not used in a manner that constitutes infringement under the current interpretation of the law. However, given the company's prior admission of book piracy, their position appears significantly weakened. The music industry's aggressive stance, exemplified by this lawsuit, suggests a coordinated effort to protect its intellectual property in the face of rapidly advancing AI technology.

The specific allegation that Mann personally torrented millions of books, and that company staff accessed further millions from pirate sites, paints a picture of deliberate and extensive data acquisition outside legal channels. Applying this same alleged methodology to the acquisition of millions of song lyrics, tied to specific datasets, presents a strong case for copyright infringement. The music publishers are not asking for a new legal interpretation; they are asking for an existing one to be applied to what they describe as a brazen campaign of IP theft.