A federal court did not rule that training an AI model on copyrighted books is illegal. It also did not rule that doing so is always legal. In Bartz v. Anthropic, the Northern District of California drew a narrower line than either headline suggests. Training on the books could plausibly qualify as fair use under US copyright law, but keeping pirated copies of those same books in order to train on them was a separate act, and that separate act was not protected. The distinction between those two questions, how the material was used versus how it was acquired, is what actually decided the case. It is also the part most coverage of the resulting 1.5 billion dollar settlement skipped past.
What the court actually found
The court considered two separate legal questions. First, whether training a large language model on copyrighted books is a transformative use that can qualify as fair use. Second, whether the way Anthropic obtained the books it trained on, downloading them from Library Genesis and Pirate Library Mirror, both known piracy repositories, was itself lawful. The answers were not the same for both questions. Training on the books could plausibly be fair use, the court found, because the model learns statistical patterns from the text rather than reproducing or distributing the books themselves. But acquiring the copies from piracy sites and retaining them in a permanent internal library was not covered by that same reasoning. Fair use can protect certain uses of a work you already hold lawfully, it does not retroactively make lawful how you obtained the copy in the first place. That second finding, about the piracy sourced library rather than the training itself, is what drove liability and shaped the settlement approved by Judge Araceli Martinez-Olguin on 20 July 2026.
Why the distinction matters beyond this one case
It is tempting to compress this into "AI training is fine" or "AI training is illegal." Neither is accurate, and treating either as a settled rule going forward misreads what happened. Fair use under US copyright law is decided fact by fact, case by case: what was copied, how much, for what purpose, and with what effect on the market for the original work. Bartz v. Anthropic answers those questions for one company, using one specific set of sources, under one specific set of facts. It is not a blanket ruling covering every AI company or every training pipeline, and other courts weighing similar questions remain free to reach different conclusions on different facts. What does generalize, more reliably than any single verdict, is the structure of the reasoning itself: a court asking two separate questions rather than one. What did you do with the material, and separately, how did you get it. A company can lose on the second question even where it might have won the first.
What this means for your own data and documents
For general counsel and compliance teams, the practical takeaway is not a legal opinion on AI training. It is a provenance question that predates AI by a long way. If your organization uses AI tools, builds on training data, or licenses content to others who might train models on it, the question worth asking is not only whether an activity is labeled training. It is where the underlying material actually came from, and whether you can show that. A model built on lawfully licensed or properly acquired material sits in a very different legal position from one built on pirated repositories, even where the training method looks identical from the outside.
The same logic runs in reverse, for material your own organization creates. Being able to show when a document existed, who created it, and that it has not been altered since is a separate, provable fact from any argument about how someone later used it. That kind of dated, verifiable record does not resolve a fair use dispute on its own, but it removes the exact ambiguity that pirated sourcing created for Anthropic in this case: ambiguity about where something came from, and whether the chain back to the original copy was ever clean. Legal and compliance teams working through IP provenance questions, for AI related risk or otherwise, can find background on how sealing and timestamping documents supports a clear chain of custody at Swiss Trust Layer's IP and legal solutions page. It will not tell you whether a given use qualifies as fair use. It is built for the adjacent, more answerable question: can you prove what you had, and when.




