Executive Summary
Thirty-five music publishing entities, Sony Music Publishing and Warner Chappell Music among them, filed a 48-page complaint against Anthropic, Dario Amodei, and Benjamin Mann on August 28, 2026, in the Northern District of California, San Jose Division. Most coverage led with the damages exposure. The part that matters to practitioners begins on page 45, in the prayer for relief.
Item (e) of nine asks the court to order an accounting of the training data, the training methods, and the known capabilities of Anthropic's AI models. It would require Anthropic to identify the lyrics and other copyrighted works it trained on, to disclose how that data was collected, copied, processed, and encoded, and to name any third parties it engaged to collect or license the data. Data lineage has been an internal documentation chore. Here it becomes an artifact a court can order into existence.
Everything below is what the plaintiffs allege. Anthropic has not yet answered the complaint, and it told Fortune this is the third lawsuit from the same lawyers recycling allegations already before the courts. Who wins is not the question here. The question is the shape of the artifact the plaintiffs are asking for.
Key Figures
The cards below hold the units of measure the complaint sets for itself. The first two are per-item statutory ceilings, the third is the volume of compositions those ceilings get multiplied against, and the last is the earlier settlement the plaintiffs invoke as their benchmark for deterrence.
Source: the complaint (Case 5:26-cv-09217, filed 2026-08-28)
$150,000
Statutory ceiling per work infringed
Where willfulness is found. 17 U.S.C. § 504(c)
$25,000
Statutory ceiling per CMI violation
§ 1203(c)(3)(B). The fourth count stands apart from infringement
Tens of thousands
Compositions listed in Exhibit B
Exhibit A, the torrenting list, runs to hundreds or more. The complaint calls both illustrative and non-exhaustive
$1.5 billion
The authors' class settlement, September 2025
The complaint calls that sum just the cost of doing business
The List on the Complaint's Last Pages
Two names lead the caption, but thirty-five sit on the plaintiff side. Twenty-four Sony Music Publishing entities and eleven Warner Chappell entities signed on as co-plaintiffs, and the caption includes the EMI catalog companies, three Hipgnosis entities, and Jobete Music, which holds Motown copyrights. The defendants are Anthropic itself plus two people: chief executive Dario Amodei and co-founder Benjamin Mann. The reason individuals are named shows up in how the counts are built.
There are four counts. Direct infringement through torrenting runs against all three defendants; contributory infringement through torrenting runs against Amodei and Mann. The remaining two name only the company. One is direct infringement through channels other than torrenting, and the other is removal and alteration of copyright management information. That last count rests on 17 U.S.C. § 1202, and it stands on its own regardless of how the infringement question comes out. Section 3 takes it up separately.
The prayer for relief runs from item (a) to item (i), nine in all. Requests for money sit alongside requests for conduct, and the damages figures the press has been calculating come out of the first two.
| Item | What is requested | Type |
|---|---|---|
| (b) | Statutory damages for willful infringement, up to $150,000 per work infringed | Money |
| (c) | Statutory damages for CMI removal and alteration, up to $25,000 per violation | Money |
| (d) | A permanent injunction against further infringement | Stop conduct |
| (e) | An accounting of the training data, training methods, and known model capabilities | Produce information |
| (f) | Destruction of infringing copies under court supervision, then a sworn report detailing how the order was carried out | Destroy and prove |
| (g)–(i) | Attorneys' fees, pre- and post-judgment interest, other relief the court deems proper | Ancillary |
▲ The prayer for relief on pages 45–47 of the complaint, grouped by type | Source: Sony Music Publishing et al. v. Anthropic PBC complaint
The money items are familiar. They are standard in any copyright suit, and with tens of thousands of compositions in Exhibit B the arithmetic that produces billions of dollars starts there. The complaint does note that Exhibits A and B are illustrative and non-exhaustive, and a footnote adds that the investigation is continuing and that the plaintiffs intend to seek leave to amend with an expanded list. Statutory damages require registration, and paragraph 57 asserts that every composition in both exhibits has been registered with the U.S. Copyright Office, clearing that threshold in advance. Any actual award still depends on proof of ownership and on how far willfulness is found.
Items (e) and (f) are the unfamiliar side. Neither asks for money. They ask for documents and for a procedure: write down what went into training, then get rid of it, then swear to how you got rid of it.
Anthropic rejects the allegations. A company spokesperson told Fortune that this is the third lawsuit from the same lawyers, recycling allegations from cases already before the courts, and said the company will defend itself robustly in court. On the training question itself the company leans on the June 2025 ruling in Bartz, where Judge William Alsup held that using copyrighted books to train a large language model that generates new text was, in purpose and character, quintessentially transformative. Counting this filing, Anthropic now faces five music copyright suits.
What the Accounting Would Have to Contain
Item (e) is a single sentence, sitting between the injunction request in (d) and the destruction order in (f).
"An order requiring Defendants to provide an accounting of the training data, training methods, and known capabilities of Anthropic's AI models, including requiring that Anthropic identify the Music Publishers' lyrics and other copyrighted works on which it has trained its AI models, and disclose the methods by which Anthropic has collected, copied, processed, and encoded this training data (including any third parties it has engaged to collect or license such data)."
The parenthetical at the end is the part that travels furthest. An accounting of your own pipeline is one problem. An accounting that reaches the vendors you engaged to collect or license data is a different one, because the records that would answer it sit partly outside your own systems: procurement contracts, dataset delivery manifests, the terms under which a labeling contractor sourced its material. Item (f) has a similar reach. It asks for destruction of infringing copies in the defendants' possession or control under the court's supervision, citing 17 U.S.C. § 503(b), followed by a sworn report setting forth in detail the manner of compliance.
The difficulty is in the four verbs. Collected asks where the material came from. Copied asks how many copies were made and where they were put. Processed asks what was added and taken away in between. Encoded asks what form the data was finally in when it entered the model. A data catalog answers none of the four. Answering them takes acquisition dates and paths, dataset versions, the rules applied at each preprocessing step, and the record of which records those rules were applied to.
Anthropic's own disclosures today are on a different scale. Paragraph 114 quotes the system cards for Claude Sonnet 5 and Claude Fable 5, both from June 2026. The description of the training data is one sentence: a proprietary mix of publicly available information from the internet, public and private datasets, and synthetic data generated by other models, plus a note that the company uses a general-purpose web crawler called ClaudeBot to obtain data from public websites. The complaint says the disclosures for earlier models were similarly sparse, and paragraph 115 alleges that Anthropic refuses to disclose the specific sources and composition of its training data. Item (e) asks for that one sentence to come back as an item-level list with a processing history attached.
Why the plaintiffs reached for this remedy is explained by a clause that recurs through the complaint. Paragraph 92 says discovery will reveal the full scope of Anthropic's copying of the publishers' works, and adds that those facts are in the possession and control of Anthropic. Paragraph 108 repeats the same clause about outputs delivered to users without the original CMI. Paragraph 100 says discovery will reveal the full extent to which the works were included in Anthropic's training data and central library. The structure is a party who cannot see inside the corpus asking the court to make the party who can hand over the list.
Items (e) and (f) close a loop between them. Write down what was in there, remove it, then prove how the removal was done. All three steps require that the assembly history of the training corpus still exist.
A prayer for relief records what a plaintiff wants from a court, not what a court has granted. The case could run for years, and there is no telling whether item (e) issues as an order in this form. Even so, the specification of the artifact is now public. The complaint has spelled out what kind of document a company would need on hand to answer a request like this one.
What Cleaning Took Out of the Text
The fourth count is aimed not at training itself but at the processing on either side of it. Copyright management information is the information attached to a work that lets a copy be identified as that work. Paragraph 179 names the titles of the compositions, the names and identifying information of the authors, and the name and identifying information of the copyright owners. Paragraph 117 counts the items the extraction algorithms allegedly stripped more finely: the copyright notice, copyright owner names, songwriter credits, performing artist name, and song title. That last item is the one that catches in practice, because it means the routine act of dropping a title field from a record falls inside what the provision covers. Section 1202 separately prohibits intentionally removing or altering this information while knowing, or having reasonable grounds to know, that doing so will induce, enable, facilitate, or conceal an infringement, and § 1203 sets statutory damages of up to $25,000 per violation. The claim stands on its own whatever the answer on fair use turns out to be.
Paragraph 102 puts the allegation this way. Anthropic cleans the text it harvests to remove material it perceives as undesirable for training, and what that cleaning removes is not the unauthorized copyrighted content, such as the publishers' lyrics, but the CMI embodied in the copied text along with other attribution and ownership information. The lyrics stay; the marks that say whose they are go. Paragraph 103 supplies the plaintiffs' reading of the motive. The expressive content generates valuable output for Anthropic's users and customers, while CMI generates no value for Anthropic and, more to the point, gives copyright owners direct evidence of infringement.
The support the plaintiffs offer is a record of tool selection. According to paragraph 105, as early as May 2021 Mann, Jared Kaplan, and other members of senior leadership weighed several content extraction tools: Newspaper, Readability, and jusText, all of them capable of separating copyright notices from the footers of webpages. In June of that year Mann and Kaplan decided against jusText because it left behind too much "useless junk," and the complaint says the junk included the critical copyright notices contained in webpage footers. An internal chat message it cites reflects dissatisfaction that jusText failed to strip a copyright owner name and a "© 2019" notice from a scraped page, while the co-founders hailed Newspaper as "a significant improvement" because it reliably removed footers, copyright owner names, and copyright notices. Those documents are not from this case. They come from the record in the earlier Concord litigation, which the complaint cites as material revealed in prior litigation.
Paragraph 106 alleges the same methods were applied to Common Crawl. The version Anthropic downloaded contained lyrics scraped from the websites of the publishers' authorized licensees, MusixMatch and LyricFind among them, and those licensees display lyrics with the corresponding CMI as their license terms require. The extraction tools, the complaint says, stripped that CMI as Anthropic curated its own training datasets. Third-party datasets appear in paragraph 96, which points out that The Pile's creators publicly disclosed in a widely published industry whitepaper that they removed copyright notices and ownership information.
The count does not stop at preprocessing. Paragraph 180 pleads two provisions of § 1202 side by side: subsection (b)(1) for intentionally removing or altering CMI, and (b)(3) for distributing works or copies of works knowing that CMI has been removed or altered. Paragraph 181 locates the removal in three places, namely the process of building the training data, the output side where models reproduce lyrics verbatim while omitting the CMI that accompanies authorized versions, and the torrenting of works from pirate libraries. Paragraph 182 folds in the case of datasets someone else had already stripped, where a party copies training data to which algorithms known to remove CMI had been applied. That is where paragraph 96 and The Pile land. It is a theory that reaches organizations that never scraped anything and trained only on public datasets.
Whether the tool change happened for the reason the plaintiffs give is a question for trial. Stripping boilerplate is standard practice in building a web corpus, and there is ample engineering justification for treating footer removal as a quality improvement. However that question resolves, one fact stays put for practitioners. What was discarded during preprocessing became evidence in a lawsuit, and a company that kept the record of it and a company that did not are not in the same position in court.
Could Your Corpus Answer Item (e) Today?
Read item (e) against your own organization and four questions come out of it. Can you identify, work by work, the copyrighted material that went into training or fine-tuning? Can you say when and by what route each piece was acquired? Can you show what preprocessing added and removed, as rules and as a record of what those rules were applied to? Can you name the third parties that collected or licensed data for you? Four yeses mean the documents already exist if item (e) ever issues as an order. Add the count from section 3 and there is a fifth question. Have you checked whether the copyright marks were already stripped from the public datasets you took in?
Item (f) adds another layer. It is not enough to have deleted particular data; the manner of deletion has to be written down and sworn to. Answering that means knowing where copies survive, not only in the originals used for training but in the intermediate files, checkpoints, and caches that fell out along the way. For an organization that has treated data deletion as a storage-cleanup task, this is the hardest item on the list.
The center of gravity of the risk is moving too. The June 2025 Bartz ruling treated the act of training as transformative use. What it left exposed was the acquisition path, and in that same case Anthropic settled for $1.5 billion after the court found it had torrented more than seven million copyrighted books. This complaint aims straight at that opening. It asserts that acquisition by torrent cannot be justified as fair use, and it builds CMI removal into a separate count so the claim does not ride on how the training question is decided. The contested question is drifting from what you trained on toward how you got it and whether you can show the process.
The same movement shows up in other music-industry litigation. In the Suno copyright case covered on this blog in July, the contested method was audio fingerprint matching to trace training tracks back to their sources. That approach infers the contents from outside. The accounting request in this complaint works the other way, pulling the records out from inside. Both routes are headed for the same thing: the assembly history of the corpus.
The case for keeping data lineage is an old one, and it has usually been made on grounds of quality control or reproducibility. This prayer for relief attaches another reason. The ledger may one day be something you are ordered to produce. The defining trait of the accounting request is that you cannot create the document on the day it is demanded.
Editor's Note
Preparing AI-Ready Data usually brings to mind getting data into a shape a model can digest. This complaint is a reminder that the shaping itself has to leave a record. Data can answer a question someone asks later only if the same ledger holds not just what went in, but what was taken out and why.
References
Primary Source
- 1.Oppenheim + Zebrak, LLP · Pryor Cashman LLP. (2026). Complaint and Demand for Jury Trial — Sony Music Publishing (US) LLC, Warner Chappell Music, Inc., et al. v. Anthropic PBC, Dario Amodei, and Benjamin Mann. U.S. District Court, N.D. Cal., San Jose Division, Case No. 5:26-cv-09217. Complaint (PDF)
Industry & News Coverage
- 2.Ingham, T. (2026, August 29). "Now Sony Music Publishing and Warner Chappell sue Anthropic in multi-billion dollar lawsuit." Music Business Worldwide
- 3.Korosec, K. (2026, August 29). "Sony Music, Warner sue Anthropic, alleging a 'brazen campaign' of intellectual property theft." TechCrunch
- 4.Hong, J. (2026, September 1). "Sony and Warner Music sue Anthropic over alleged theft of 'tens of thousands' of songs." Fortune
- 5.Chen, J. (2026, August 29). "Sony and Warner sue Anthropic for 'blatant violation' of copyright law." Engadget