US DOJ Backs OpenAI in Copyright Case, Warns AI Training Restrictions Threaten National Security

What Happened

On 1 September 2026 the U.S. Department of Justice filed a Statement of Interest of the United States in the Southern District of New York in In re: OpenAI, Inc. Copyright Infringement Litigation, MDL No. 25-md-3143 (SHS) (OTW). The 19 page filing relates to all consolidated matters and is signed by Associate Attorney General Stanley E. Woodward Jr., Assistant Attorney General Brett Shumate and Senior Counsel Michael Weisbuch. It became public on 2 September 2026 through Reuters, the Associated Press and MLex.

The position is stated directly: the United States has a strong interest in the court rejecting any argument that training large language models on copyrighted texts violates copyright law.

The scope of the government’s position is narrower than it might seem, but not narrow enough to bypass the question of outputs entirely. The government does limit its transformative-use argument to the training stage, and in footnote 2 clarifies that it does not assert the defendants’ conduct was authorized by or performed for the benefit of the United States within the meaning of 28 U.S.C. § 1498. Still, the department does address outputs when it argues that, absent substantial similarity to copyrighted works, they cannot cause market harm to rights holders, one of the fair use factors concerning market effect.

Procedural Status

The filing is made under 28 U.S.C. § 517, which allows the Department of Justice to attend to the interests of the United States in any pending suit, with no time limit and no need for the court’s leave. A statement of interest does not bind the court and is not a source of law: it is an executive position placed on the docket immediately before summary judgment, where Judge Sidney H. Stein will decide fair use on the merits. Formally the United States does not appear for a party. In substance the filing tracks OpenAI’s defence.

Training as Transformative Use

The United States quotes Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007, 1021 (N.D. Cal. 2025): copying protected text as part of training an LLM is a use of a different kind or character that is transformative, spectacularly so. It adds Kadrey v. Meta Platforms, Inc., 788 F. Supp. 3d 1026, 1044 (N.D. Cal. 2025).

The anchor is Google LLC v. Oracle America, Inc., 593 U.S. 1 (2021): if precise copying of code was transformative because it let programmers put existing skills to use in a new system, then copying a work to teach a model to recognise relationships between data is at least as transformative.

An important qualification: the analysis follows Andy Warhol Foundation v. Goldsmith, 598 U.S. 508 (2023) and proceeds use by use. The government accepts that certain output uses may not be transformative where a model reconstructs and disseminates an original work, but insists this should not bear on the training use.

Market Dilution Called “Deeply Flawed”

This is the core of the filing. The United States attacks the market dilution theory in Kadrey, 788 F. Supp. 3d at 1052-57, under which generated works erode the market for originals even absent reproduction. The government notes the theory was adopted without the benefit of briefing and without a single supporting authority.

Two objections:

  • The court improperly collapsed training and outputs into a single, continuous use and applied an overly broad, genre-level understanding of substitution. Training is its own use and involves no substitution: it reveals nothing to the public.
  • The theory overlooks substantial similarity, the most basic element of infringement. Outputs that compete with an article but reproduce none of its protected elements are not substitutes merely because they sit in the same genre, which is an uncopyrightable idea or method of expression.

The government’s illustration is Joan Didion, who as a teenager typed out Hemingway’s stories to learn how the sentences worked. On the Kadrey logic she would have been liable to Hemingway for every piece she published, because training and creation would be one use and her work competed with his.

The filing separately criticises the Register of Copyrights’ position in Part 3 of the Copyright Office report (May 2025), arguing it warrants no deference after Loper Bright Enterprises v. Raimondo, 603 U.S. 369 (2024).

The section closes on separation of powers: whether this technology warrants rewriting copyright principles is a policy question for the legislature.

The Oligopoly Argument

The least conventional part of the filing rests on competition policy. An erroneous fair use ruling would hamper competition because only the largest technology companies might have the capital necessary to pay licensing fees. In the filing’s words, it is not in the public’s interest for those companies to gain an oligopoly over LLM training as a result of licensing barriers to entry that function primarily as large subsidies for old mainstream media companies. Footnote 13 qualifies this: the United States takes no position on whether a licensing regime would be feasible, and notes that publishers already license access to paywalled and proprietary content.

Regulatory Backdrop and Reaction

The position rests on the Executive Orders of 23 January 2025 and 2 June 2026 and on the National Policy Framework for Artificial Intelligence of March 2026, which records that training AI models on copyrighted material does not, in and of itself, violate copyright law. In the same days Commerce Secretary Howard Lutnick urged G20 countries to allow AI training on creators’ work.

The American reasoning does not carry over to Europe automatically: the EU relies on the text and data mining exceptions in article 4(3) Directive 2019/790 (DSM) with a rightsholder opt out, plus the copyright policy duty under article 53 of the AI Act.

The New York Times said the administration is siding with trillion dollar AI companies at the expense of the American creators whose work they stole. The Intercept, itself a plaintiff, and The Hollywood Reporter described the filing as direct executive support for the defence.

Next Steps and Why It Matters

Cross motions for summary judgment follow. Under Judge Stein’s amended schedule, reply briefs are due by 6 November 2026 and no hearing date is set. The ruling will be the first full federal decision on training models on journalistic content.

  • The dispute shifts from the training stage to the output stage. If the court accepts the split between the two uses, litigation will turn on specific outputs and their similarity to the work.
  • Rejecting market dilution removes the argument that required no proof of similarity and raises the evidentiary burden on plaintiffs.
  • Lawful acquisition of data remains a separate risk: in Bartz the court accepted training as fair use but assessed retention of a pirated library separately. The Statement of Interest does not address it.
  • The oligopoly argument cuts both ways and confirms that licensing pay walled and proprietary content remains ordinary practice whatever the outcome.

Practical Recommendations

  • Separate the training layer from the output layer in documentation, policies and contracts, as distinct processes with distinct legal bases.
  • Document lawful acquisition of every dataset: scraping vendor agreements, source terms of use, licences.
  • Deploy and document output side controls that prevent verbatim reproduction of protected text.
  • For rightsholders, build the record at output level: prompts, model responses, comparative similarity analysis.
  • State permission or prohibition of training use expressly in licence and user terms, and back it with technical means: robots.txt, machine-readable metadata and other technical means of expressing a TDM reservation of rights, API access terms.
  • Do not carry the US reasoning into European or UK transactions without adapting it to the TDM regime and article 53 of the AI Act. Under the AI Act, obligations attach to placing a model on the EU market, regardless of where the training took place. This means that even if a model’s training was lawful in the US, an EU court will not necessarily reach the same conclusion, as happened in GEMA v. Suno. We previously covered GEMA v. OpenAI, which also raised the issue of training.

We Are Ready to Help

The Arbitration & IT Disputes practice at REVERA Law Group is ready to advise on risks arising from the use of protected content for model training, audit training data sources and supplier contracts, prepare licensing documentation for rightsholders and developers, and act in disputes with model developers and major online platforms.

Write to us










    Send request