Executive Summary

On September 1, 2026, the United States filed twenty pages in the Southern District of New York. It is not a ruling and not a party's brief. It is a statement of interest, and it is addressed not to the New York Times case alone but to every copyright suit against OpenAI now gathered in one consolidated docket. The government's argument compresses into a single sentence: copying copyrighted text in order to train a language model is fair use. The following day in Chapel Hill, the same government, chairing the meeting, saw a G20 innovation ministers' consensus statement adopted. The phrase "training data" does not appear in that document once.

Open the filing and you find the government splitting language-model development into three stages, acquisition, training, and output, writing that each may raise distinct questions of copyright law, and then committing itself to the middle one. The stage it sets aside already carries a price. Anthropic agreed to pay $1.5 billion in a case where training itself was held to be fair use and the liability attached to where the books had been obtained. Europe turned the same question into a filing requirement, and since last month that filing has carried a fine.

This report puts one question to all three documents: what does each demand as evidence? The three tracks differ in their conclusions, in who decides, and in when the decision lands. They converge on one prior requirement: a record of which works were obtained from where, and by what route. That record is needed whichever way the rulings go, which makes provenance less a bet on a regulatory direction than the one option that loses under none of them.

Four numbers give the scale of the three tracks before the argument starts: the price already attached to acquisition, the ceiling on the fines now circulating in Europe, how fully the summaries filed under those fines were actually completed, and the size of the litigation this filing lands in.

~$3,000

Per work, Anthropic settlement

$1.5B ÷ roughly 482,000 works · the price already set on acquisition

3% of revenue

EU ceiling for disclosure breaches

€15M or 3% of global turnover, whichever is higher · Article 101, live since 2026-08-02

3 of 7

Providers who answered in prose

As of 2026-08-07 · the form is compulsory, its contents are not

19

Actions pending in the MDL

MDL 25-md-3143 · JPML statistics, report date 2026-09-01

1

Twenty Pages Filed in the Southern District of New York

The document is captioned Statement of Interest of the United States and appears under 28 U.S.C. § 517. The case is In re: OpenAI, Inc. Copyright Infringement Litigation, 25-md-3143 (SHS)(OTW), and the header line reads "This Document Relates To: All Matters." That line puts every case gathered in the consolidated docket inside the filing's reach. Footnote 11 makes it explicit: the government says it refers only to the New York Times and OpenAI to simplify, and that its legal arguments "apply similarly to all parties in this litigation and the related cases, including book authors and publishers."

The signature block carries Associate Attorney General Stanley E. Woodward, Jr. and Assistant Attorney General Brett Shumate of the Civil Division, with Senior Counsel Michael Weisbuch signing. The filing date is September 1, 2026. Because the coverage and the Chapel Hill G20 statement landed in the same news cycle the following day, the three events read as simultaneous. A day separates the court filing from the international communiqué, and the two documents were doing different work.

For its interest the government cites two executive orders and the President's National Policy Framework for Artificial Intelligence, then adds a national security argument. It quotes the Government Accountability Office on AI capabilities and national security, and continues:

"Rules of law that make it significantly more difficult to develop a robust AI industry in the United States therefore threaten national security and give a competitive advantage to foreign adversaries who are not so encumbered."

Statement of Interest, page 3

The competition argument is the stranger one. Mandatory licensing, the filing reasons, would leave only the largest technology companies able to pay, hardening training into an oligopoly, and the fees collected along the way would function "primarily as large subsidies for old mainstream media companies." That sentence relocates copyright enforcement from a question about protecting creators to a question about barriers to entry. Footnote 13 then steps back: the United States "takes no position on whether a licensing regime would be financially or logistically feasible."

Open the policy document the government leans on and the emphasis runs in more than one direction. The filing quotes the National Policy Framework saying that "American creators, publishers, and innovators should be protected from AI-generated outputs that infringe their protected content, without undermining lawful innovation and free expression." That sentence puts protection at the output stage. The filing goes on to record that the President has encouraged Congress to consider "enabling licensing frameworks or collective rights systems for rights holders," with a parenthetical that such legislation should not address when or whether licensing is required, alongside a federal framework for digital replicas and continued monitoring of copyright law. And in the very next sentence the filing cites the same document again for the proposition that the "training of AI models on copyrighted material," in and of itself, "does not violate copyright laws."

One policy document, then, holds both a sentence that releases training and a sentence that holds on to outputs. Twenty pages go to arguing the first. The second passes by in a single line and drops into the footnotes. Which stage to defend was this filing's first editorial decision, and the evidence for that is inside the authority the government chose for itself.

The other side answered the next day. "The Administration is siding with a handful of trillion-dollar AI companies at the expense of the countless American creators whose work they stole," Times spokesman Graham James told Deadline. He added: "Both AI and creators can thrive – AI companies simply need to pay fairly for the content that makes their products possible, as copyright law requires," and said the administration's proposal "to let companies take that content without permission or compensation would undermine the sustainability of the human-created content that a healthy society depends on, and which AI needs to function." We were not able to confirm any formal statement from creators' or copyright organizations about this filing. The one public response on record so far comes from a plaintiff.

1.1What This Document Actually Weighs

The filing carries less weight than its caption suggests. It is not a ruling, it does not bind the court, and the United States is not a party. In footnote 2 the government disclaims any contention that the conduct at issue was authorized or consented to by the government under 28 U.S.C. § 1498. The statute itself never uses the words "statement of interest." Section 517 says only that the Department of Justice may "attend to the interests of the United States in a suit pending in a court of the United States," and courts have read that broadly enough to accept non-party filings as a matter of practice. Footnote 1 cites Gil v. Winn Dixie Stores (242 F. Supp. 3d 1315, 1317) for the point that the statute "contains no time limitation and does not require the Court's leave." Filing freely and being adopted are separate things. Christine Bartholomew has argued that § 517 filings can become a channel for executive overreach.

The most conspicuous passage in the filing is footnote 17. The government notes that the Register of Copyrights appeared to endorse a market-dilution theory in a report, then writes "Her understanding does not warrant deference" and cites Loper Bright (603 U.S. 369, 2024), the decision that withdrew judicial deference from agency readings of the law. A § 517 statement of interest is not a document that receives deference either, and copyright is not the Justice Department's statute to administer. The government swings the no-deference principle at the Copyright Office while standing exactly where the same principle applies to its own filing.

Footnote 17 carries one more thing. The government identifies the Register of Copyrights as an official "who is currently challenging her removal", in the body of the footnote and without hedging, and then calls the reasoning in her report "threadbare". A person's employment dispute and the weight of the report that person wrote sit inside one sentence. How that dispute is proceeding falls outside this article and we take no view on how it ends. That an argument for withholding deference from copyright interpretation arrived in court while the official channel for that interpretation was itself unsettled is worth putting on the record.

1.2The Conflict-of-Interest Question

The reporting establishes this much. On July 2, 2026 the Financial Times, citing two people familiar with the matter, reported that OpenAI had been discussing with the administration a plan to hand the U.S. government a 5% stake; CNBC and Axios followed. The Financial Times characterized these as "early conversations." Against the $852 billion valuation set at the March 2026 round, 5% works out to roughly $42.6 billion. OpenAI declined to comment and the White House did not respond. Nothing has been agreed, and the filing does not mention the discussions.

From here on it is commentary. Writing in Above the Law on September 3, 2026, Joe Patrice pointed to the absence of any disclosure of those stake discussions in the filing and called the intervention selective. We do not carry that criticism over as fact. What is established is that the discussions were reported, that the filing does not record them, and that the silence has drawn criticism.

Another data-related front in the same litigation has already been covered here. The fight over preservation of chatbot conversation logs is set out in New York Times Seeks Sanctions Over OpenAI's Deleted Chatbot Logs.

2

What the Government Defended Was the Training Stage

Section II.A of the filing describes model development in three stages. At the acquisition stage, also called collection or pre-training, a developer "first collect[s] data, including plaintiffs' works." Then the developer trains the model by feeding that data through it. At the output stage, the model responds to user queries. None of this framing is the government's invention. It is lifted from the court hearing this case, in New York Times Co. v. Microsoft Corp. (No. 1:23-cv-11195-SHS-OTW, ECF No. 514 at 6–8, April 4, 2025). And in the sentence that follows, the government narrows its own range.

"Each stage may present distinct questions of copyright law. The United States focuses on the question whether the use of copyrighted works at the training stage—by copying works in order to feed data into the model as learning material—constitutes fair use."

Statement of Interest, page 6

Three boxes, taken from the court, and the middle one selected for defense. Nowhere in these twenty pages does the government defend acquisition. It says nothing about it at all. Output goes into footnote 15, sorted away as a separate matter: a tiny sliver of anomalous reconstructive outputs, the filing argues, would not support a remedy that cuts off LLM output uses generally, let alone training. Footnote 16 adds that training by itself makes no copied work accessible to the public, and that even at the output stage at most a very small fraction of outputs would expose protected aspects of those works.

Filing §II.A's Three Stages — What the Government Actually Defended Acquisition Training Output Not defended Already priced Anthropic settlement $1.5B Defended Fair use argument All twenty pages of argument Pushed to footnote Unresolved · contested nv-recall 4.0–95.8% Reconstructed from the filing's §II.A three-stage framing | Original Pebblous diagram
▲ The three stages of language-model development and what the U.S. government's filing actually defended — acquisition and output are left blank; twenty pages argue training alone

The difficulty is that the first box, the one left empty, already has a number in it. Bartz v. Anthropic (787 F. Supp. 3d 1007) put it there. Training won in that case. Judge Alsup called the use of the works for training "transformative—spectacularly so," and the government's own filing cites that page, 1021, as support. The money attached somewhere else: to roughly seven million books pulled from shadow libraries such as LibGen and PiLiMi and kept. Divide the $1.5 billion settlement by the roughly 482,000 works covered and each work comes to about $3,000. Final approval came on July 20, 2026, from Judge Martínez-Olguín, after Judge Alsup had retired.

Compressing that case into "Anthropic lost on fair use" would collapse this article's whole argument. Training won; obtaining the books lost. The box the U.S. government defended and the box that cost Anthropic $1.5 billion are two different boxes. The sequence is set out in Anthropic's Copyright Settlement Draws a Legal Line on Data Acquisition.

2.1Two Ways to Handle Acquisition

The market has already produced two answers to the same problem. By LLM Pulse's tally, current as of September 2026, OpenAI leads with 24 disclosed publisher licensing deals, the largest being a five-year agreement with News Corp worth more than $250 million. In the same tally Anthropic has zero disclosed publisher deals. It settled the acquisition question afterward instead, for $1.5 billion. One company buys the acquisition stage up front; the other pays for it later. The running total of 91 deals we used in AI No Longer Buys Data. It Rents It. comes from a different compiler with a different denominator, so it does not belong in the same table as the 24.

Footnote 13 concedes that the licensing market exists. Regardless of how fair use comes out, it says, "both mainstream and independent publishers could enter (and have entered) into licensing agreements to provide developers with specialized access to real-time, pay-walled, proprietary, and other content and information." In the government's argument that sentence works as evidence that licensing need not be compelled. Read from inside an organization it says something else. The route of buying acquisition through contract is already open, and having taken that route is provable only through the contracts and the records.

3

The Output Stage, Left in a Footnote

The stage the government pushed into a footnote comes with one flat assertion attached. It is the government's contention rather than a settled fact.

"Training an LLM, in and of itself, does not make any copied works accessible to the public. Even at the distinct output stage, at most a very small fraction of outputs would make protected aspects of those works accessible."

Statement of Interest, page 12, footnote 16

Early this year a study measured that assertion directly. Extracting books from production language models (arXiv:2601.02671), released on January 6, 2026 by Ahmed Ahmed, A. Feder Cooper, Sanmi Koyejo and Percy Liang, went after four commercially deployed models, the kind that ship with safety systems in front of them. The procedure runs in two phases: an initial probe for extraction feasibility, sometimes using a Best-of-N jailbreak, then iterative continuation prompts to pull the book out. Success is scored as nv-recall, a block-based approximation of longest common substring.

Model Jailbreak required nv-recall
Gemini 2.5 Pro No 76.8%
Grok 3 No 70.3%
Claude 3.7 Sonnet Yes (Best-of-N) 95.8%
GPT-4.1 Yes; roughly 20× more attempts, then refuses 4.0%

Source: Ahmed et al., arXiv:2601.02671 (2026-01-06), abstract. The first two figures come from the Phase 1 probe against Harry Potter and the Sorcerer's Stone. The authors state that they ran different per-LLM experimental configurations, so these four numbers are not a like-for-like ranking.

Putting the two sets of numbers head to head would be a misreading. The government's "very small fraction" refers to how often ordinary use produces infringing output. The paper's nv-recall measures how much of a specific targeted book can be recovered under attack. The denominators differ, so the paper cannot be written up as a rebuttal of the government's claim. The two sentences answer different questions. What the research does establish is the authors' own conclusion: even with model-level and system-level safeguards, extraction of in-copyright training data remains a risk for production LLMs. The output stage is not a closed question.

The last row belongs to the model of the defendant in this litigation. Of the four it held out hardest. It demanded roughly twenty times more attempts and then refused to continue. Resistance at the output stage varies by company, and the variation is a product of design rather than luck. It is also worth stating that the targeted work, Harry Potter and the Sorcerer's Stone, circulates in many editions and is unusually well represented in text. The result does not transfer unchanged to an arbitrary news article.

3.1Using One Ruling in Both Directions

Read the filing to the end and Kadrey v. Meta Platforms (788 F. Supp. 3d 1026) turns up twice, put to opposite uses. Page 1044 supplies support: because language models produce tools that edit email, translate paragraphs and write skits, training is highly transformative. The market-dilution discussion at 1052–57 becomes the target. The government labels it "Contrary dicta," calls that court's application of the fourth fair-use factor "deeply flawed," and argues that the Kadrey court collapsed training and outputs into a single continuous use, did so without the benefit of briefing, and itself acknowledged that the indirect-substitution theory had never made a difference in a case before. We were not able to open the cited pages of the Kadrey opinion directly. What is written above records only what the filing says.

The government's most memorable rebuttal never went into a footnote at all; it sits in the body of the filing. As a teenager, Joan Didion "would type out" Ernest Hemingway's "stories to learn how the sentences worked," and later named him the greatest influence on her writing. By the Kadrey court's logic, the filing asks, Didion should have incurred liability to Hemingway every time she published a piece. The source is her Paris Review interview.

4

At Chapel Hill, the G20 Sent the Question Home

A day later, on September 2, the G20 innovation ministers adopted their consensus statement in Chapel Hill, North Carolina, with the United States in the chair. The document is built on six pillars: pro-innovation policy frameworks, technology for opportunity and prosperity, skilled technical workforce development, intellectual property policies for artificial intelligence, AI for standards and standards for AI, and industrial innovation and investment in supply chains. Copyright in training data has exactly one place to live here, the fourth pillar.

"…domestic and international legal frameworks and doctrines, such as prior and express consent and applicable limitations and exceptions to copyright, have played a critical role in balancing the interests of creators and innovators in some members' jurisdictions, and that their application to AI-related activities remains appropriately resolved through each member's established legal processes."

G20 Innovation Ministerial Consensus Statement, Pillar 4

That single sentence steps back three times. It lists the European approach of prior and express consent alongside the American approach of limitations and exceptions without endorsing either. It then confines even that role to "some members' jurisdictions," which turns a universal principle into an observation about certain countries. And it hands the question of application to each member's own legal processes. The preamble runs the same way: the ministers call on members to develop their own policies and preserve national sovereignty in the governance of emerging technologies.

Stepping back is not the same as saying nothing. The pillar first records the "essential role that copyright protections play in safeguarding and supporting the creative works of authors, artists, innovators, and other rightsholders," and acknowledges that the interaction between copyright law and AI raises complex legal and policy questions across jurisdictions. The sentence that closes the pillar treats enforcement of IP rules and trade-secret protections, "fair and appropriate remuneration" and remedies against misappropriation as important. The vocabulary overlaps with what the plaintiff's side had said a day earlier about paying fairly. This sentence too ends with a tail: "in accordance with applicable legal frameworks." Acknowledge the principle, return the execution to national law. The structure repeats through the pillar.

A word-level pass over the full statement sharpens the picture. Training data appears zero times, text and data mining zero, dataset zero, provenance zero, fair use zero. Consent occurs once, in the sentence quoted above. All three occurrences of copyright fall inside a single Pillar 4 paragraph, one in each of three consecutive sentences. The four transparency-family terms are spread across public-sector adoption conditions, standards, and supply chains. The one in Pillar 6 concerns sharing supply-chain information and comes marked "on a voluntary basis."

The G20 Chapel Hill Consensus Statement — Word Counts Every training-data keyword: zero. Only general-principle words appear training data 0 text and data mining 0 dataset 0 provenance 0 fair use 0 consent 1 copyright 3 transparency 4 Full-text word count · one of the four "transparency" hits sits in Pillar 2's institutional-trust context | Original Pebblous diagram
▲ In the full text of the G20 Chapel Hill consensus statement, all five training-data keywords return zero hits — only general terms like copyright and transparency appear at all

We also checked the Carolina Principles released the same day. It has three sections, on advancing discovery and strengthening technology development, accelerating validation and commercialization, and enabling technology adoption. Copyright and consent appear zero times each. Intellectual property appears exactly once, in an item about sharing intermediate results from discontinued projects, inside the qualifier "while respecting applicable legal frameworks." It is not a provision that governs IP. It is a promise not to trespass on it.

Within a day Washington produced two documents. In court it took training data head on and argued a substantive position. In the international statement it never used the phrase and returned the judgment to member states. We would not call that a contradiction. One is a litigation filing about the interpretation of domestic law; the other is a diplomatic text premised on respect for sovereignty. The observation that the two documents point in different directions is as far as we take it, and the rest belongs to the reader.

One note on wording: the White House release says ministers attended, not that they signed. Reporting this statement in terms of how many countries signed would add something the source does not say. For anyone running an organization, the practical residue is straightforward. Standards will continue to diverge by jurisdiction for the foreseeable future, and international consensus is not what will resolve that divergence.

5

In Europe the Same Question Carries a Fine

A phrase that circulates in the industry needs correcting first. "The EU starts enforcing training-data transparency on August 2" runs two things together. The obligations on general-purpose AI providers, technical documentation, a public summary of training content, and a copyright policy, have applied since August 2, 2025. What switched on for the first time on August 2, 2026 was not the rule but the teeth. The Commission's AI Office gained the power to fine breaches, under Article 101, at €15 million or 3% of global turnover, whichever is higher. Article 50 transparency duties and the market-surveillance enforcement machinery started running the same day.

Which switches flipped and which were pushed to 2027 is laid out in On August 2, the EU AI Act's switches aren't the ones you think and On August 2, Chatbots and Deepfakes Have to Come Clean, and the larger design, in which Europe placed this question in advance through the text-and-data-mining exception and machine-readable opt-outs, is in Europe's Top Court Asks First: What Is an LLM Allowed to Learn From?. What remains is how those filings were actually completed.

5.1How the Filed Summaries Were Filled In

The researcher Kieran Maynard compared 11 summaries from 7 providers field by field against the template, and that comparison is written up in Big Tech Left Half of the EU's Training-Data Summaries Blank. Google, Meta, Microsoft, OpenAI and the open-source contingent filled in the template's quantitative fields. The other three, Anthropic, Mistral and xAI, replaced those fields with narrative phrasing such as "a proprietary mix of publicly available information and licensed data." Even after enforcement powers took effect, as of August 7, 2026 there were zero confirmed fines or investigations. What has happened since then we have not verified.

Three out of seven is a statement about the limit of the regime rather than its failure. The form is already compulsory. Failing to file exposes a provider to sanction, and the summary has to be refreshed every six months. What the rules do not compel is what goes in the boxes. A single narrative sentence satisfies the paperwork while telling a reader nothing about which works came from where. That is exactly where organizations diverge in front of the same template. Only the ones holding work-level acquisition records can complete the quantitative fields.

OpenAI left a signal of its own. Its statement on EU AI Act compliance, published on July 31, 2026, two days before enforcement powers arrived, covered its safety framework, watermarking partnerships and cybersecurity cooperation in detail while omitting the copyright chapter entirely, meaning both the training-content summary and the copyright compliance policy. That does not mean no summary was filed. It does show that watching which item drops out of a company's own account of its compliance is another way of seeing that meeting the form and disclosing the substance are two different achievements.

6

One Corpus, Three Kinds of Proof

Lay the three documents over one organization's corpus and the first thing visible is how differently they ask. They differ in kind, in who does the judging, in when the judgment falls, and in what happens on failure. The table below separates those differences item by item.

Item U.S. courts G20 Chapel Hill EU AI Act
Nature of the document A defense raised in litigation Non-binding principles A legal obligation
Who decides The presiding judge No one; left to members The Commission's AI Office
When it is judged After the fact Continuously
What you produce Evidence filed in court A published summary
Cost of failure Damages None €15M or 3% of revenue
Minimum record demanded Acquisition route per work Varies by jurisdiction Composition of training content

The three cells in that last row are worded differently and start from the same place. To raise fair use you have to show that this work was obtained lawfully. To answer a jurisdiction that requires consent you have to show that consent was given. To file a summary in Europe you have to write down what came from where. Three demands drawing on one record in three formats. Without the record, none of the three tracks can even be started.

6.1Where the Industry Actually Stands

The Data Provenance Initiative, whose contributors include MIT and Harvard Law School, traced and audited more than 1,800 text datasets on popular online repositories. Licenses were left unspecified in over 70% of cases. Where a license was stated, more than half were wrong. Narrowing to Hugging Face, 66% were categorized under a use category different from what the original creators had stated, and the misclassification generally ran more permissive than reality. These figures describe a 2023–2024 sample of public fine-tuning datasets rather than an estimate for full training corpora. Even with that limit attached, the conclusion holds. An organization that took license labels at face value cannot prove what it trained on. What verification actually costs in hours and dollars is measured separately in The Real Invoice for Free Data.

6.2What to Record, and at What Unit

Decomposing what all three tracks ask in common gives the following, and the point is less the contents of the list than the unit. A corpus-level note saying "a mix of public web and licensed data" may clear the European template while proving nothing at all in court.

  • Source URL or an identifier for where it was obtained
  • Date and time of acquisition
  • Method of acquisition: purchase, license agreement, crawl, donation, or internal generation
  • Rights basis documents: contracts, terms-of-use snapshots, license notices
  • The license text as stated, and when it was checked
  • Redistribution history: what path the data traveled to reach you

This is precisely the information that disappears the moment data is re-downloaded. Reconstruction reaches as far as crawl logs, purchase records and contracts survive, and no further. That training data is already being asked for the way ledgers are asked for is visible in Music Publishers Ask a Court to Order an Accounting of Claude's Training Data.

6.3If the Court Sides With the Government

Suppose the opposite. Judge Stein adopts the government's position outright, and training is settled as fair use across the board. Does the need for acquisition records shrink? It does not.

First, what the government defended was the training stage alone, and acquisition stays empty. That empty box is where Anthropic's $1.5 billion landed. Second, Europe's disclosure duty runs independently of any American ruling and already carries fines. Third, once the G20 has returned the judgment to each member's legal processes, standards stay split by jurisdiction, and in a jurisdiction that requires consent the record is the qualification. Any one of the three tracks surviving on its own is enough to keep work-level acquisition records necessary.

7

Why Pebblous Is Watching This

Acquisition looks like a legal problem, and the place it actually gets handled is the data pipeline. Which document arrived when, from where, and by what method is knowable only at the moment the data enters the pipeline; miss that moment and everything afterward is inference. That is why the source and route metadata produced when a corpus is diagnosed turns out to be the shared evidence across all three tracks. The definition of AI-Ready Data has carried the condition from the start that you must be able to prove the data was permissible to train on.

Seen that way, acquisition route stops being supplementary information and becomes a quality attribute, because the same sentence in the same document is usable or not depending on where it came from. Which means data of unknown origin is not simply data with a label missing. It is better treated as defective. Reading the 70% of public-repository datasets with no license label as a quality metric rather than a compliance headline follows from the same logic.

Put into practice, there is something you can settle while the ruling is still pending: what unit of record to keep. Attaching the six fields from the previous section at the level of individual works is the starting point, and starting honestly means accepting that data already used in training can be reconstructed only as far as crawl logs and purchase records reach.

Editor's Note. Pebblous watches this subject because what the three regulatory tracks have begun to demand in common is whether the origin and route of a dataset can be proven through diagnosis, and that question overlaps with the one DataClinic has been aimed at. If a record is what is required whichever way the regulatory outcome falls, then an organization that keeps acquisition route as a standing byproduct of its pipeline has finished preparing before the outcome arrives. Pebblous's AI-Ready Data pipeline and DataClinic sit adjacent to where that diagnosis happens. Neither substitutes for legal advice or for licensing services.

The twenty pages filed on September 1 are not a ruling and do not bind the court. Even so, the most useful information the document leaves behind is not what the government argued but what it left blank. The blank box was one that already had a price on it, Europe has converted the same box into a filing, and the G20 has sent the judgment back to national law. When the same record is what all three directions ask for first, there is no reason to wait for the day a verdict arrives.

R

References

This article was written against the primary documents listed below, obtained directly. The full text of the statement of interest and both G20 documents were downloaded, and every quotation and word count was verified against the source. The cited pages of the Kadrey opinion and of the Copyright Office's Part 3 report could not be opened directly; everything drawn from them is reported here only as what the statement of interest says.

Primary documents

  • 1.Statement of Interest of the United States, In re: OpenAI, Inc. Copyright Infringement Litigation, No. 25-md-3143 (SHS)(OTW), ECF No. 316 (S.D.N.Y. filed Sept. 1, 2026).
  • 2.The White House, "G20 Innovation Ministerial Concludes with Consensus Statement" (Sept. 2, 2026). Link
  • 3.The Carolina Principles for Emerging Technologies, G20 Chapel Hill (Sept. 2, 2026).
  • 4.JPML, MDL Statistics Report — Distribution of Pending MDL Dockets by Actions Pending, Report Date Sept. 1, 2026. Link
  • 5.JPML, MDL No. 3143 Transfer Order (Apr. 3, 2025). Link

Research

  • 6.Ahmed Ahmed, A. Feder Cooper, Sanmi Koyejo & Percy Liang, "Extracting books from production language models," arXiv:2601.02671 (Jan. 6, 2026). Link
  • 7.The Data Provenance Initiative, "A Large-Scale Audit of Dataset Licensing and Attribution in AI," arXiv:2310.16787 / Nature Machine Intelligence (2024). Link
  • 8.Christine P. Bartholomew, "The Dark Side of Antitrust Statements of Interest," Journal of Corporation Law (2024).

Policy & case law

  • 9.U.S. Copyright Office, Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version, May 2025). Link
  • 10.Loper Bright Enterprises v. Raimondo, 603 U.S. 369 (2024).
  • 11.Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007 (N.D. Cal. 2025) — cited via the statement of interest.
  • 12.Kadrey v. Meta Platforms, Inc., 788 F. Supp. 3d 1026 (N.D. Cal. 2025) — cited via the statement of interest.
  • 13.Gil v. Winn Dixie Stores, Inc., 242 F. Supp. 3d 1315 (S.D. Fla. 2017).

Press & industry trackers

  • 14.Financial Times, "OpenAI discusses giving US government 5% stake" (July 2, 2026), with follow-up coverage by CNBC and Axios.
  • 15.Joe Patrice, Above the Law (Sept. 3, 2026); IPWatchdog, "DOJ Sides with OpenAI…" (Sept. 3, 2026).
  • 16.LLM Pulse, "Every AI Content Licensing Deal, Mapped (2023–2026)." Link
  • 17.Deadline, "NY Times Rips Trump's DOJ For Backing AI Companies In Class Action Suit" (Sept. 3, 2026) — statement by Times spokesman Graham James. Link

Related Pebblous publications

  • 18.Pebblous, "Anthropic's Copyright Settlement Draws a Legal Line on Data Acquisition." Link
  • 19.Pebblous, "On August 2, the EU AI Act's switches aren't the ones you think." Link
  • 20.Pebblous, "Big Tech Left Half of the EU's Training-Data Summaries Blank." Link
  • 21.Pebblous, "Europe's Top Court Asks First: What Is an LLM Allowed to Learn From?" Link
  • 22.Pebblous, "AI No Longer Buys Data. It Rents It." Link