AI and Copyright

Artificial intelligence puts pressure on copyright from two directions at once, and almost every confused argument about it comes from mixing them together. The first question is about inputs: whether copying protected works to train a model infringes, or whether it is fair use. The second is about outputs: whether what the model produces can be owned at all, and by whom. These are governed by different doctrines, they are being decided in different forums, and they can come out in opposite directions in the same case.

As of mid-2026 the output question has a reasonably firm answer and the input question does not. Keep them apart and the landscape becomes legible.

Why the two questions are genuinely separate

Training a model requires copying. Works are downloaded, stored, converted, and processed, and each of those steps implicates the reproduction right in 17 U.S.C. § 106. That is an infringement question, met by the fair use defense in § 107.

Ownership of the output is a different inquiry entirely. It runs through § 102(a), which extends protection to “original works of authorship,” and through the constitutional and statutory meaning of “author.” Nothing about winning or losing the training argument tells you who owns the result. A developer could train lawfully and still produce output nobody can copyright.

The input question: is training fair use?

Three 2025 district court decisions frame the whole debate, and they do not line up neatly.

In Bartz v. Anthropic, Judge William Alsup held in June 2025 that training a large language model on books was transformative enough to be fair use. But he refused to extend that shelter to the millions of copies Anthropic had downloaded from pirate libraries such as Library Genesis and kept in a permanent internal collection. Acquiring the copy and training on the copy were separate acts, and only the second was excused. The piracy claim never reached a jury: Anthropic settled for roughly $1.5 billion covering about 500,000 works, at approximately $3,000 per work, the largest copyright settlement in U.S. history. It received preliminary approval in September 2025, and after a fairness hearing on May 14, 2026 the court took final approval under submission, where it remains.

Days later, in Kadrey v. Meta, Judge Vince Chhabria granted summary judgment to Meta on the training claim, calling the use highly transformative. His reasoning was a warning shot rather than a victory lap. He held that the plaintiffs lost because they failed to build a record on market dilution, the theory that flooding a market with machine-generated substitutes harms authors even when no output reproduces any particular book. He wrote plainly that on a better record, many such plaintiffs should win.

Thomson Reuters v. Ross Intelligence cuts the other way. In February 2025, Judge Stephanos Bibas rejected fair use where Ross used Westlaw headnotes to train a legal research tool that competed head-on with Westlaw. The distinguishing feature was that Ross’s product was not generative. It retrieved judicial opinions rather than producing new expression, which made the transformative argument much weaker and the market harm much more direct.

No federal appeals court has decided any of this. The Third Circuit heard oral argument in Ross on June 11, 2026, which makes it the first appellate court positioned to speak. Until it does, every confident claim about whether AI training is fair use is a prediction, not a statement of law.

The pattern emerging from the input cases

Three variables are doing most of the work across these decisions:

  • How you got the copy. Lawful acquisition and pirated acquisition are being treated as different acts with different consequences, and this is currently the sharpest dividing line in the case law. Both Bartz and Kadrey left torrenting and distribution claims alive after resolving the training claim.
  • Whether the system generates or retrieves. Generative systems have fared better on the transformative-purpose factor than tools that substitute for the original’s function.
  • Market harm, factor four. This is where the cases turn. The Supreme Court’s decision in Andy Warhol Foundation v. Goldsmith (2023) narrowed how much work “transformative” can do when a use competes commercially with the original, and market dilution is the live theory plaintiffs are now building records around.

The U.S. Copyright Office reached a similar posture in Part 3 of its AI report, released in pre-publication form in May 2025: some training is fair use, some is not, and the answer depends on the source of the data, the nature of the model, and the effect on the market. It declined to endorse a categorical rule in either direction.

The output question: human authorship is required

Here the law is settled, at least at the threshold. Copyright protects works of human authorship. Machine output, standing alone, is not protectable by anyone.

Stephen Thaler tested this directly, seeking registration for an image he said his “Creativity Machine” authored autonomously, with himself as owner. The Copyright Office refused. The D.C. Circuit affirmed in March 2025, holding that the Copyright Act requires human authorship as a matter of statutory construction. The Supreme Court denied certiorari in March 2026, leaving the rule in place.

The Copyright Office’s Part 2 report on copyrightability, issued January 2025, applies the same principle to the ordinary case where a human uses a tool. Its central conclusion is that prompts alone do not confer authorship, because prompts influence output without dictating it. Effort does not substitute for authorship. The “sweat of the brow” doctrine died in Feist Publications v. Rural Telephone (1991), and iterating through hundreds of prompts does not revive it.

What you can actually own

The registration of Zarya of the Dawn shows how the line gets drawn in practice. Kris Kashtanova made a comic book with Midjourney-generated images. The Office registered the work, but only in part. Kashtanova’s human-written text and the creative selection and arrangement of the images were protected. The images themselves were not, because Midjourney generated them unpredictably rather than executing a conception Kashtanova had formed.

That leaves three reliable sources of protection in AI-assisted work: original human expression you contribute, modifications to machine output substantial enough to qualify on their own, and the compilation or arrangement of generated elements. Registration requires disclosing more than trivial AI-generated content and disclaiming it. The human layer survives; the machine layer is public domain, free for anyone to use.

What remains unsettled

Quite a lot, and honesty about that is more useful than false precision. Whether training is fair use awaits appellate review. Whether market dilution is a cognizable harm under factor four is being litigated now. Whether outputs that regurgitate training data create liability, and for whom, is at the center of New York Times v. OpenAI, now in discovery with trial expected around late 2026 or 2027. Music has largely detoured around the courts: Universal settled with Udio in October 2025 and licensed its catalog, Warner settled with both Udio and Suno, while other labels are still litigating for a precedent. Getty v. Stability produced a UK ruling in November 2025 that Stability largely won, on claims Getty had narrowed before judgment, and it tells you little about U.S. law. Disney v. Midjourney is on a schedule that pushes any merits ruling toward 2027.

Expect movement. Anything you read about AI and copyright, including this page, has a shelf life measured in months.

Frequently asked questions

Can AI-generated art or writing be copyrighted? Not the machine-generated part. U.S. copyright requires a human author, a rule the D.C. Circuit affirmed in Thaler v. Perlmutter (2025) and the Supreme Court declined to review in March 2026. Prompts alone do not make you the author of the output, because they influence the result without dictating it. What you can protect is your own human contribution: original text you wrote, edits substantial enough to qualify on their own, and the creative selection and arrangement of AI-generated elements.

Is training an AI model on copyrighted works fair use? Unresolved, and the answer so far depends heavily on facts. Two 2025 district courts, Bartz v. Anthropic and Kadrey v. Meta, held that training a generative model on books was transformative and qualified as fair use, while Thomson Reuters v. Ross rejected fair use for a non-generative legal research tool that competed directly with the source. No federal appeals court has ruled. The Third Circuit heard argument in Ross in June 2026 and its decision will be the first appellate word.

Does using AI to help write or draw something destroy my copyright? No. Assistive use is fine, and the Copyright Office has been explicit that AI tools in the workflow do not disqualify a work. Protection simply does not extend to the portions the machine generated. When you register a work with more than a trivial amount of AI-generated material, you must disclose it and disclaim that material, which leaves your human authorship protected and the rest in the public domain.

Who is liable if an AI output copies someone’s work? Potentially the developer, the user, or both. Copyright infringement under 17 U.S.C. § 501 is a strict liability tort, so not knowing the output copied something is not a defense. A user who publishes an infringing output is the one distributing it, while the developer faces direct claims over the copies made during training and secondary claims for contributing to user infringement. Courts have not yet allocated this cleanly.

Authorities and sources

Going further: AI and Intellectual Property, the practical guide .

This page is general legal information, not legal advice, and it does not create an attorney-client relationship.

The cases behind this
AI & Copyright

Bartz v. Anthropic: Transformative Training, Unforgivable Acquisition

Judge Alsup held that training a large language model on books is 'exceedingly transformative' fair use, while refusing to extend that blessing to the pirated library that fed it. The $1.5 billion settlement that followed shows where the real exposure lies.

June 24, 2026
AI & Copyright

Kadrey v. Meta: A Fair-Use Win That Reads Like a Plaintiffs' Brief

Two days after Bartz, Judge Chhabria also found AI training to be fair use, but went out of his way to say the result reflected a failure of advocacy, not a vindication of the practice. His 'market dilution' theory is the doctrine to watch.

June 22, 2026
AI & Copyright

Thomson Reuters v. Ross: The First Refusal of Fair Use in the AI Era

Before the generative-AI rulings, a Delaware court rejected fair use for using copyrighted material to build an AI legal-research tool, and pointedly distinguished the software cases the technology industry had relied upon. Its reach is narrower than its reputation.

June 19, 2026
Practical Guides
More in Copyright