# Bartz v. Anthropic: Transformative Training, Unforgivable Acquisition

> Judge Alsup held that training a large language model on books is 'exceedingly transformative' fair use, while refusing to extend that blessing to the pirated library that fed it. The $1.5 billion settlement that followed shows where the real exposure lies.

Topic: Copyright  |  Author: Lidiia Levitska  |  Source: Intellectual Property Law (outsideipcounsel.com)
Canonical: https://outsideipcounsel.com/blog/bartz-v-anthropic-ai-training-fair-use/


The first generation of generative-AI copyright decisions has arrived not as a single doctrinal pronouncement but as a set of carefully partitioned holdings, and *Bartz v. Anthropic PBC*, No. 3:24-cv-05417-WHA (N.D. Cal. June 23, 2025), is the most instructive of them. Senior Judge William Alsup did not answer the binary question the parties pressed on him (is training fair use, yes or no) but instead disaggregated the defendant's conduct into three distinct acts and subjected each to its own analysis under 17 U.S.C. § 107. The result is a decision that hands the AI industry its most important win to date on the core training question while simultaneously exposing the practice that has proven most expensive: the acquisition of the underlying corpus.

## At a glance

- **Case:** *Bartz v. Anthropic PBC*, No. 3:24-cv-05417-WHA (N.D. Cal.)
- **Decided:** June 23, 2025 (order on partial summary judgment), Senior Judge William H. Alsup
- **Holding:** Training Claude on lawfully acquired books is fair use; destructive scanning of purchased print copies is fair use; downloading and retaining millions of pirated books is not fair use
- **Status:** Fair-use ruling stands; the piracy claim resolved through a reported $1.5 billion settlement, with final approval and distribution pending as of mid-2026

## The doctrinal frame: § 107 and transformative use

Fair use is an affirmative defense, evaluated through the four non-exclusive factors codified at 17 U.S.C. § 107: the purpose and character of the use (including whether it is commercial and whether it is "transformative"), the nature of the copyrighted work, the amount and substantiality of what was taken, and the effect on the potential market for the work. Since *Campbell v. Acuff-Rose Music* the first factor has carried the conceptual weight of the inquiry, and the Supreme Court's recent decision in *Andy Warhol Foundation v. Goldsmith* recalibrated it toward a comparison of purposes: does the secondary use merely supersede the original, or does it serve a sufficiently different end?

What makes *Bartz* methodologically important is that Judge Alsup refused to run a single fair-use analysis across the defendant's entire course of conduct. He recognized that "using books to build an AI" is not one act but several, each with a different relationship to the copyrighted works, and that § 107 must be applied to each.

## Three acts, three answers

**Training.** The use of lawfully acquired books to train Anthropic's Claude models was, in the court's words, "exceedingly transformative." The animating analogy is that a model ingests text in order to learn the statistical structure of language and produce new output, much as a human reads in order to learn to write. The transformative character of that use under the first factor outweighed the creative nature of the works under the second: the books were expressive, core-copyright material, but that did not rescue the plaintiffs once the purpose of the use was found to diverge so sharply from the purpose of the originals.

**Format conversion.** Anthropic had also destructively scanned print books it lawfully purchased, converting them into a digital, searchable internal format. The court held this, too, was fair use: a change of medium for copies the defendant already owned, with no new copies distributed to the public, is the kind of internal format-shifting that the doctrine tolerates.

**Acquisition by piracy.** The third act broke the other way, and decisively. Anthropic had downloaded and retained more than seven million pirated books from shadow libraries such as LibGen to assemble a permanent central library. That wholesale copying from infringing sources was not excused by the transformative ends to which the copies were eventually put; in the court's words, "piracy of otherwise available copies is inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and immediately discarded." Nor could Anthropic cure the problem by buying the books afterward. The closing line of the order is the one that matters for compliance: that Anthropic "later bought a copy of a book it earlier stole off the internet will not absolve it of liability for the theft but it may affect the extent of statutory damages." The legality of the *use* does not cure the illegality of the *acquisition*. That claim was left for trial.

## Why the partition matters

The analytical move is more consequential than the headline. By refusing to treat "AI training" as a monolith, Judge Alsup decoupled the fair-use question from the provenance question. A developer may possess an entirely defensible transformative-use theory and still face crippling liability for how it sourced its data. That reorients the litigation risk: the contested issue in the next wave of cases is less likely to be the abstract permissibility of training than the discoverable facts of corpus assembly: what was downloaded, from where, and whether the developer ever held a lawful copy of each work.

It also clarifies what *Bartz* does not decide. This was a summary-judgment ruling on a developed record before a single district judge; it binds no court outside the case. Its persuasive force on the training question is genuine, but it travels alongside a piracy holding that should temper any reading of the decision as a clean victory for AI developers. The opinion is better understood as a map of where risk lives than as a safe harbor.

## The settlement as epilogue, and the lesson in the number

The piracy claim was set for trial in December 2025. It never reached a jury. The parties settled for $1.5 billion, reported as the largest copyright settlement in United States history, covering a works list of 482,460 titles. The district court granted preliminary approval on September 25, 2025, and the case was reassigned to Judge Araceli Martínez-Olguín; at a fairness hearing on May 14, 2026, the court took the matter under submission rather than granting immediate final approval, and distribution remained pending as of mid-July 2026. Class counsel reported that 440,490 of the 482,460 works had been claimed, a 91.3 percent claims rate, as of mid-April 2026.

The number is the lesson. A transformative-use holding of the first order did not insulate the defendant from ten-figure exposure, because the liability lived in the acquisition, not the use. The economics are worth stating plainly: a defensible model-training program can coexist with catastrophic copyright liability if the training data was assembled from infringing sources.

## Open questions

Several issues *Bartz* leaves unresolved will shape the doctrine. First, the "transformative training" holding rests on a record about how these models ingest and abstract text; a different record (for instance, one showing memorization and verbatim regurgitation of training works in output) could alter the first- and fourth-factor analysis. Second, Judge Alsup did address the fourth-factor market-dilution theory but rejected it outright, reasoning that the authors' complaint was "no different than it would be if they complained that training schoolchildren to write well would result in an explosion of competing works," which is "not the kind of competitive or creative displacement that concerns the Copyright Act." Two days later, in *Kadrey v. Meta*, Judge Chhabria took the opposite view of that same theory, calling market dilution the most promising route for authors even as he granted summary judgment to Meta on the record before him. Which judge is right remains open. Third, the piracy holding raises but does not resolve how a developer "launders" a tainted corpus: whether subsequent licensing or purchase can mitigate, rather than cure, liability remains open.

## Implications for businesses and developers

- **Data provenance is the principal control.** Maintain auditable records establishing lawful acquisition of every work in a training corpus. The transformative nature of training will not excuse infringing sourcing.
- **Distinguish acquisition from use in compliance design.** Internal format-shifting of lawfully owned copies sits on firmer ground than ingestion of materials obtained from unauthorized repositories.
- **Treat shadow-library sourcing as a strict-liability risk.** *Bartz* indicates that downloading from pirate sources is independently actionable regardless of downstream purpose.

## Frequently asked questions

**Did the court rule that AI training is legal?** It held that, on this record, training Claude on lawfully acquired books was fair use. That is a single district court's ruling, not a categorical or binding rule, and it was paired with a holding that pirating the training data is not protected.

**Why did Anthropic pay $1.5 billion if it won on fair use?** Because it won only on the *use* of lawfully acquired works. The separate claim (that it had downloaded and retained millions of pirated books) was not protected by fair use and was resolved by settlement before trial.

**What is the practical takeaway for companies building AI?** How you obtain training data may matter more than how you use it. Lawful, documented sourcing is the central compliance obligation the decision identifies.

## Authorities and sources

- Order on summary judgment, *Bartz v. Anthropic PBC*, No. 3:24-cv-05417-WHA (N.D. Cal. June 23, 2025): [opinion PDF (Copyright Alliance)](https://copyrightalliance.org/wp-content/uploads/2025/06/Bartz-v.-Anthropic-Order.pdf); [docket (CourtListener)](https://www.courtlistener.com/docket/69058235/bartz-v-anthropic-pbc/).
- Analysis: [Authors Alliance, "Anthropic Wins on Fair Use for Training its LLMs; Loses on Building a 'Central Library' of Pirated Books"](https://www.authorsalliance.org/2025/06/24/anthropic-wins-on-fair-use-for-training-its-llms-loses-on-building-a-central-library-of-pirated-books/).
- Settlement: [Authors Guild, "What Authors Need to Know About the Anthropic Settlement"](https://authorsguild.org/advocacy/artificial-intelligence/what-authors-need-to-know-about-the-anthropic-settlement/); [Publishers Weekly on the May 2026 fairness hearing](https://www.publishersweekly.com/pw/by-topic/digital/copyright/article/100438-little-drama-at-anthropic-s-settlement-hearing.html).

