Sony Music Publishing and Warner Chappell Music have filed an Anthropic music lawsuit alleging that the AI company used tens of thousands of copyrighted songs to develop Claude without permission. The complaint, filed in California on 28 August 2026, names Anthropic and two of its co-founders, Dario Amodei and Benjamin Mann. Its claims are allegations that the court has not yet decided.
The publishers say Anthropic copied lyrics and musical compositions from unauthorised sources, removed copyright information and enabled Claude to reproduce protected material. Anthropic says the case recycles allegations already before the courts and has promised to defend itself. The dispute could test not only whether AI training is transformative, but whether the way training data was obtained changes the legal outcome.
What Sony and Warner allege
The 48-page federal complaint was filed in the US District Court for the Northern District of California, San Jose Division. The plaintiffs include Sony Music Publishing, Warner Chappell and numerous affiliated companies that administer songwriting catalogues.
They allege that Anthropic used copyrighted musical compositions as training inputs for Claude and retained material in a central library. The complaint says Exhibit B identifies tens of thousands of works and gives examples including Ain’t No Mountain High Enough, All I Want for Christmas Is You, Eye of the Tiger, Here Comes Santa Claus and Paper Rings.
The action is not simply a claim that a model learned from music. It alleges specific acts of copying, including torrenting large collections, scraping authorised lyrics services and using third-party datasets containing copyrighted text. It also says physical books were bought, cut apart and scanned to create digital copies.
Those allegations place the case inside the wider fight over AI training data lawsuits. Courts are being asked to examine several stages separately: obtaining a work, storing it, using it for training and producing outputs that may reproduce protected expression.
Why the source of training material matters
Anthropic has previously won an important but limited US ruling over training on books. In June 2025, Judge William Alsup held in Bartz v Anthropic that using lawfully obtained books to train large language models was fair use. He also held that converting purchased print books into digital copies for the company’s own library was fair use.
However, the court’s order refused to treat allegedly pirated library copies as protected merely because they were later used for training. The distinction was between the transformative purpose of training and the separate act of acquiring and retaining unauthorised copies.
Sony and Warner are trying to bring that distinction into music. Their complaint alleges that Anthropic downloaded material from sources including LibGen and PiLiMi, scraped lyrics from services whose licences did not authorise AI training and relied on datasets such as Books3. Anthropic has not admitted those allegations in this case.
Even if a court accepts that some training uses can be fair, it could still ask whether the underlying copies were obtained legally. That is why a broad statement such as “AI training is fair use” does not settle every training-data dispute. The source, purpose, retention and resulting outputs can raise different questions.
Lyrics, outputs and copyright information
The publishers also allege that Claude produced verbatim or near-verbatim lyrics and derivative versions of their songs. Output claims matter because copyright generally protects original expression, not a bare fact or general style. A system that recalls substantial protected passages presents a different issue from one that generates a new work after learning statistical patterns.
The complaint adds claims involving copyright management information, often shortened to CMI. This can include details identifying a work, its author or rights owner. The publishers allege that identifying information was removed or altered during data processing and in model outputs.
Anthropic will be able to contest the examples, their legal significance and the connection between any training copy and a particular output. At this stage, the complaint states one side’s case. It is evidence of what the publishers allege, not proof that the alleged conduct occurred.
The output issue is especially sensitive in music because lyrics can be short, memorable and commercially licensed across several formats. LiveAIWire’s earlier examination of how generative AI is rewriting the music industry showed why questions of consent, attribution and payment are becoming central to the technology’s adoption.
Anthropic says AI training is fair use
An Anthropic spokesperson told Reuters that this was the third lawsuit from the same lawyers and that it recycled allegations from cases already before the courts. The company said it would defend itself robustly and pointed to the earlier judgment recognising AI training as fair use.
That response addresses the company’s broad legal position but does not resolve the new complaint’s factual claims. The publishers represent different rights, identify a large set of musical compositions and include allegations about lyrics, data sources and copyright information. The court will need to decide which claims survive and what evidence each side can establish.
Fair use in the United States is a case-specific defence. It considers the purpose and character of the use, the nature of the copyrighted work, the amount used and the effect on the potential market. The US Copyright Office’s report on generative AI training likewise concluded that licensing questions depend on facts and market conditions rather than a single rule covering every model and dataset.
What the publishers want from the court
Sony and Warner seek a judgment that Anthropic infringed their copyrights, along with an injunction restricting the challenged conduct. They also ask for damages and profits, or statutory damages where available.
The complaint cites potential statutory damages of up to US$150,000 for each work if infringement is proved and found wilful. It also seeks remedies for alleged violations involving copyright management information, including statutory awards allowed by law. Those figures are maximum amounts requested under legal provisions, not a calculated judgment or an amount Anthropic currently owes.
With tens of thousands of compositions listed, the theoretical exposure sounds enormous. Actual damages would depend on which works qualify, whether infringement is established, the court’s interpretation of the claims and any later settlement. Large headline arithmetic should not be mistaken for a forecast of the final result.
Why creators and AI companies are watching
Music publishers license compositions for recordings, broadcasts, films, games, performances and digital services. If AI training becomes another established licensing market, model developers may face higher costs and more detailed obligations about provenance. If broad training uses are held to be fair, rights holders may focus more heavily on unauthorised acquisition and reproducing outputs.
The creative sector is already adapting to that uncertainty. Questions about training data sit beside concerns over synthetic competition, attribution and the bargaining power of individual creators. LiveAIWire’s reporting on generative AI in creative industries found that adoption was moving faster than a stable consensus on rights and payment.
AI companies also need dependable rules. Training runs require large datasets and substantial investment, while unresolved provenance can create legal risk years after data is collected. A model developer may therefore care as much about documented acquisition and filtering as about the ultimate scope of fair use.
What happens next in the Anthropic music lawsuit
Anthropic is expected to answer the complaint or challenge parts of it through motions. The court may then narrow the claims before discovery examines data sources, internal decisions, model behaviour and the publishers’ ownership records. That process can take months or years, and many copyright disputes settle before trial.
The immediate lesson is narrower than either side’s preferred slogan. A prior ruling favourable to AI training did not declare every route to a training library lawful. Equally, a new complaint containing tens of thousands of titles does not establish infringement by itself.
The Sony and Warner case will be watched because it connects three unsettled issues in one action: how copyrighted data was acquired, whether training can qualify as fair use and when generated output crosses into unlawful reproduction. Until the court tests the evidence, the most accurate description is also the simplest: the publishers have made sweeping allegations, and Anthropic has said it will fight them.
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.
