Anthropic Used Torrent Site Music to Develop Claude AI Model
Anthropic faces a major copyright lawsuit from two major music publishers over allegations that it used BitTorrent to obtain unlicensed sheet music and songbooks for Claude’s training. The case hinges on whether downloading from pirate sites constitutes willful infringement, a distinction that could expose the company and named executives to statutory damages reaching $150,000 per work.
- Sony Music Publishing and Warner Chappell Music allege Anthropic downloaded roughly 5 million books from Library Genesis in June 2021 and 2 million from Pirate Library Mirror in July 2022
- Publishers also claim Anthropic scraped lyrics from Musixmatch and LyricFind, which pay for display rights to the material
- Complaint names Dario Amodei and Benjamin Mann as individuals rather than employees, potentially subjecting them to personal liability and deposition
- $150,000 Maximum statutory damages per song for willful infringement as sought by publishers
- 7 million Total books downloaded from pirate libraries between June 2021 and July 2022
- 2023 Year Anthropic previously defeated these same publishers’ attempt to block Claude training
Sony Music Publishing and Warner Chappell Music filed suit against Anthropic on Friday, alleging the company systematically obtained unlicensed music through torrent downloads and web scraping to train its Claude AI model. The complaint identifies two specific data hauls from shadow libraries: approximately 5 million books from Library Genesis downloaded in June 2021, and roughly 2 million from Pirate Library Mirror in July 2022. The publishers contend that sheet music and songbooks embedded in those collections were then used to train Claude, and they further allege that Anthropic scraped song lyrics directly from Musixmatch and LyricFind, services that pay licensing fees for distribution rights.
The lawsuit emerges against a backdrop of intensifying legal disputes over AI training data. Multiple music publishers, authors, and rights holders have challenged various AI companies over their use of copyrighted material, arguing that training artificial intelligence systems on protected works without permission or compensation constitutes large-scale infringement. The music industry has been particularly vocal about such practices, given the sector’s historical reliance on licensing revenue and performing rights organizations that ensure creators receive compensation when their work is used.
Torrenting carries legal risks that book scanning did not
A federal judge has already established a critical precedent distinguishing between these two acquisition methods. In prior litigation, the same court found that Anthropic’s practice of purchasing physical books and scanning them fell within fair use protections, allowing that program to survive legal challenge. However, the court rejected downloads from piracy sites as “straightforward piracy but at massive scale,” leading Anthropic to settle that dispute and cease the practice.
BitTorrent introduces an additional legal complication beyond mere downloading. The torrent protocol requires uploading while downloading, meaning every copy obtained is simultaneously shared with other users. The complaint’s first two of four counts rest heavily on this distribution mechanism, treating the transmission as infringement in itself rather than relying solely on training use. This distinction matters because settlements resolve corporate liability, while individuals face depositions and potential personal judgment.
Legal experts have noted that the willfulness question represents a critical juncture in AI copyright litigation. If a court determines that Anthropic knowingly violated copyright law, statutory damages multiply substantially, potentially reaching the maximum amounts the publishers seek. Evidence that company leadership understood the legal risks yet proceeded anyway would strengthen claims of willful infringement and expose executives to enhanced penalties.
Dario amodei and benjamin mann named as individuals, not employees
The complaint names Anthropic’s Chief Executive Dario Amodei and co-founder Benjamin Mann personally rather than solely in their corporate capacity. This designation opens the door to personal liability and testimony under oath, a significantly higher bar than corporate settlement negotiations typically face. Companies can settle lawsuits; individuals may be compelled to explain their knowledge and decisions during discovery.
The publishers claim hundreds of songs sat within the pirate libraries and request up to $150,000 per work found to have been willfully infringed, with the broader training claim encompassing tens of thousands of compositions. Notably, Anthropic defeated an earlier bid by these same publishers in 2023 to block Claude’s training over lyric scraping claims, agreeing instead to implement output guardrails that limit the model’s reproduction of protected lyrics.
The prior settlement attempt suggests these publishers have been monitoring Anthropic’s practices and remain unconvinced about the company’s compliance with copyright law. The decision to pursue this lawsuit rather than negotiate further indicates the publishers view the evidence of systematic piracy as sufficiently compelling to warrant federal litigation, despite the costs and uncertainties of court proceedings.
Discovery will determine whether Music reached claude’s Training pipeline
The case now hinges on a central factual question: whether the downloaded music and lyrics actually entered Claude’s training data or remained unused. Anthropic has acknowledged torrenting books but has never conceded that music was included in those files, positioning this lawsuit as a test of whether the company’s knowledge and intent can be proven in discovery.
Technical analysis will likely play a significant role in these proceedings. Forensic examination of training datasets, metadata analysis of the downloaded files, and documentation of data processing procedures could reveal whether sheet music and lyrics were intentionally retained or filtered out during preparation. Communications among employees regarding data sourcing decisions will also face scrutiny, as courts assess whether personnel at Anthropic understood they were acquiring content from unauthorized sources.
Anthropic has not publicly commented on the suit. The resolution will depend on documentary evidence and testimony revealing whether the sheet music and lyrics were deliberately filtered into training or excluded, and whether executives knowingly participated in the downloads from pirate sites. The outcome could significantly shape how AI companies approach training data acquisition and whether individual executives face personal consequences for corporate data practices.
BlockWest is a news publication. Nothing here is investment advice. Read our disclaimer and editorial policy.
