Tremblay v. OpenAI: The Pleading Bar for Output-Based Infringement
A federal court trimmed authors' claims against OpenAI, holding that calling every ChatGPT output an infringing derivative fails to plead similarity.
The first wave of lawsuits pitting authors against large language model developers forced courts to confront a novel question with old tools: when a chatbot is trained on copyrighted books, what exactly is the infringement, and how must a plaintiff plead it? Tremblay v. OpenAI, Inc., No. 3:23-cv-03223 (N.D. Cal.), gave one of the earliest substantive answers. On February 12, 2024, Judge Araceli Martinez-Olguin of the United States District Court for the Northern District of California granted in part and denied in part OpenAI’s motion to dismiss, cutting back several of the authors’ theories while leaving the core of the case intact.
The decision matters less for any final ruling on liability, which it did not reach, than for the pleading discipline it imposed. The court held that a plaintiff cannot simply declare that every output of a generative model is an infringing derivative of the plaintiff’s work. That shortcut, the court reasoned, skips the essential element of substantial similarity. The ruling set an early marker that later suits, including those brought by other authors and news organizations, have had to navigate.
At a glance
- Case: Tremblay v. OpenAI, Inc., No. 3:23-cv-03223 (N.D. Cal. Feb. 12, 2024).
- Decided: February 12, 2024; order by Judge Araceli Martinez-Olguin granting in part and denying in part the motion to dismiss.
- Holding: Authors’ vicarious-infringement and DMCA Section 1202 theories were dismissed, with leave to amend, because the complaint did not allege that any specific ChatGPT output was substantially similar to the plaintiffs’ books.
- Status: Ongoing as of July 2026. The direct-infringement theory premised on training copies survived, and on April 3, 2025 the Judicial Panel on Multidistrict Litigation transferred Tremblay and eleven related actions to the Southern District of New York for coordinated pretrial proceedings before Judge Sidney H. Stein, as In re: OpenAI, Inc., Copyright Infringement Litigation, MDL No. 3143.
The two theories of AI copyright liability
To follow the ruling, separate the two distinct infringement theories that recur across the AI copyright cases. The first is an input or training theory: the developer allegedly made unauthorized copies of the plaintiffs’ books when it assembled and processed the training corpus. That is a direct-infringement theory grounded in the reproduction right, and it does not depend on what the model later produces.
The second is an output theory: the model’s generations are themselves infringing because they reproduce or adapt the plaintiffs’ protected expression. An output theory sounds in the reproduction and derivative-work rights, but it carries copyright law’s familiar proof requirement. To show that an accused work infringes, a plaintiff must establish that it is substantially similar to protected elements of the original. A conclusory assertion that outputs are derivative, untethered to any comparison of a particular output against a particular book, does not satisfy that requirement. Tremblay is fundamentally about the second theory and the discipline it demands.
The facts and posture
The plaintiffs, authors Paul Tremblay and Mona Awad, with related plaintiffs in consolidated proceedings, alleged that OpenAI copied their books to train its large language models and that ChatGPT could produce summaries and derivative content based on those books. They pleaded a bundle of claims: direct copyright infringement, vicarious copyright infringement, violation of the DMCA’s copyright-management-information provisions under 17 U.S.C. Section 1202(b), negligence, unjust enrichment, and unfair competition under California law. OpenAI moved to dismiss most of the complaint.
The court’s order sorted the claims. It dismissed the vicarious infringement claim, the DMCA claim, the negligence claim, and the unjust enrichment claim, granting leave to amend and setting a March 13, 2024 deadline for an amended complaint. It allowed the California unfair competition claim to proceed under the statute’s unfair prong, while the unlawful and fraudulent prongs fell along with the DMCA claim they rested on. The direct-infringement theory tied to the alleged copying of the books during training was not the subject of the dismissal, leaving that central claim in the case.
Why the output claims failed
On vicarious infringement, the court held that the plaintiffs had not adequately alleged an underlying direct infringement by the outputs. The complaint asserted that ChatGPT generates infringing derivative works, but it did not “explain what the outputs entail or allege that any particular output is substantially similar – or similar at all – to their books.” Without that factual predicate, there was no direct infringement to which vicarious liability could attach, so the derivative-liability claim could not stand. The court dismissed it with leave to amend, inviting the plaintiffs to plead specific similar outputs if they could.
The DMCA Section 1202 claim failed for related reasons. The plaintiffs alleged that OpenAI removed copyright-management information when it copied their books into the training data. But Section 1202(b) requires a showing that the defendant knew or had reasonable grounds to know that removing the information would induce, enable, facilitate, or conceal an infringement. The court found the plaintiffs had not shown how omitting management information in training copies would give OpenAI reasonable grounds to know that ChatGPT’s outputs would facilitate infringement, and it noted that the plaintiffs had not alleged OpenAI distributed their books or identical copies of them, only supposed derivatives described without detail. The claim was dismissed.
What the ruling changed for AI litigation
Tremblay did not decide whether training on copyrighted works is lawful, nor whether any output infringes. What it did was insist that plaintiffs bringing output-based theories meet copyright’s ordinary substantial-similarity pleading standard rather than substitute a categorical assertion that all model outputs are derivative. That is a meaningful constraint. It channels output claims toward concrete evidence, requiring plaintiffs to produce and compare specific generations against specific protected passages.
The order also foreshadowed the burden that has shaped the broader docket. Later cases, including other authors’ suits and the high-profile disputes involving news publishers, have had to plead their infringement theories with more precision or lean on the training-copy theory that Tremblay left standing. The decision thus helped separate the durable claims, direct copying for training, from the more speculative ones, blanket output infringement, that require particularized proof to survive.
Open questions
- Can output claims be repleaded successfully? Tremblay dismissed with leave to amend. Whether plaintiffs can produce specific outputs substantially similar to their books, and how courts will assess such comparisons, remains a live and evolving question as of July 2026.
- Is training itself infringement or fair use? The surviving direct-infringement theory squarely raises whether copying works to train a model is actionable or protected fair use, an issue the dismissal order did not resolve.
- How does Section 1202 apply to training data? The court’s treatment of the double-knowledge requirement leaves unsettled how CMI-removal claims fare when the removal occurs in a training pipeline rather than in a distributed copy.
Implications for creators and businesses
- Plead specific outputs, not slogans. Rights holders pursuing output-based claims against AI systems should be prepared to identify actual generations and compare them, element by element, to protected expression. Categorical assertions will not survive a motion to dismiss.
- The training-copy theory is the durable one. Direct infringement based on the making of training copies remains the sturdiest claim, and the fair-use battle over that conduct is where much of the war will be fought.
- AI developers should document their pipelines. Because Section 1202 and fair-use defenses turn on what was copied, stored, and stripped, developers benefit from clear records of how training data is acquired and processed.
- Expect claims to be sorted, not swept aside. Tremblay shows courts trimming weak theories while preserving stronger ones, so both sides should focus resources on the claims most likely to endure rather than litigating everything at once.
Frequently asked questions
What did the court dismiss in Tremblay v. OpenAI? On February 12, 2024, the court dismissed the authors’ vicarious copyright infringement, DMCA Section 1202, negligence, and unjust enrichment claims, with leave to amend. It let a state unfair competition claim proceed under the unfair prong. The direct infringement theory based on unauthorized training copies was not part of the dismissal.
Why did the vicarious infringement claim fail? Because the plaintiffs did not allege that any particular ChatGPT output was substantially similar to their books. The court held that asserting every output is an infringing derivative work, without showing similarity between a specific output and a specific book, does not state a claim.
Does Tremblay mean AI outputs can never infringe? No. It holds only that a plaintiff must plead facts showing a specific output is substantially similar to a specific protected work, rather than relying on a blanket assertion. Output-based claims remain viable if pleaded with that particularity, and the training-copy theory was untouched.
Authorities and sources
- Tremblay v. OpenAI, Inc., No. 3:23-cv-03223 (N.D. Cal. Feb. 12, 2024) (FindLaw)
- In re: OpenAI, Inc., Copyright Infringement Litigation, MDL No. 3143, transfer order (J.P.M.L. Apr. 3, 2025) (govinfo)
- Loeb & Loeb, “Tremblay v. OpenAI, Inc.”
- Knowing Machines, “Tremblay v. OpenAI” explainer
- 17 U.S.C. § 1202, integrity of copyright management information (Cornell LII)