Everything we know about how Anthropic's Project Panama scanned millions of books
Court records detail millions of pirated downloads, millions of purchased print books, a destructive scanning workflow and the 482,460-title settlement that followed.
By Ryan Merket · Published · Updated
Primary source: X - Ole Lehmann
Why it matters
The case gives AI developers a working legal boundary: model training and internal scanning of purchased books may qualify as fair use, while pirated acquisition can create billion-dollar liability.

Anthropic's Project Panama was an internal effort to build a large digital book library by purchasing print copies, cutting off their bindings, scanning the pages into searchable PDFs and discarding the paper originals. Court records exposed the operation alongside a separate book-acquisition track that ultimately cost the AI developer $1.5 billion: downloads of more than 7 million pirated copies.
A federal judge approved the settlement on July 20th, resolving claims tied to 482,460 books obtained from Library Genesis and Pirate Library Mirror. The ruling did not reject Anthropic's physical scanning method. Instead, it drew a legal line between digitizing books the company had bought and retaining books acquired without payment.
Here is what the court record shows about why Dario and Daniela Amodei's company created Project Panama, how the operation worked, the numbers behind Anthropic's book collection and what the settlement does and does not cover.
The numbers behind Anthropic's book collection
Anthropic's two acquisition tracks operated at different scales and carried different legal risks:
- Co-founder Ben Mann (@8enmann) downloaded at least 5 million books from LibGen in June 2021.
- Anthropic downloaded at least 2 million more books from Pirate Library Mirror in July 2022.
- The company had previously acquired 196,640 titles from Books3.
- Judge William Alsup found that Anthropic pirated more than 7 million book copies in total.
- Anthropic later spent millions of dollars purchasing millions of physical books, many of them used, for scanning.
- The settlement covers 482,460 books obtained from LibGen and Pirate Library Mirror.
- Authors and publishers had claimed interests in 440,490 of those works, or 91.3% of the settlement list, as of April 16th.
- The $1.5 billion settlement represents an estimated payment of roughly $3,000 per covered book before final allocation adjustments.
The more than 7 million pirated copies and the 482,460 books covered by the settlement are not interchangeable figures. The larger number describes Anthropic's pirate-library downloads. The settlement list identifies the works included in the resolved claims.
Why Anthropic wanted a book library
Anthropic's founders concluded early in the company's development that books were among the most cost-effective sources for building a competitive large language model. Anthropic, founded in 2021 and led by siblings Dario and Daniela Amodei, publicly described its mission as creating AI systems that were "steerable, interpretable, and robust," according to its Series A announcement.
Court records show that satisfying the data requirements of frontier-model development became a separate operational challenge. Anthropic initially assembled a central digital library from Books3, LibGen and Pirate Library Mirror. The company retained that library for general research, including books it had decided not to use for model training, according to a June 23rd, 2025 court order.
That distinction mattered in court. Anthropic's exposure did not rest solely on whether particular books entered a training corpus. Alsup treated the creation and retention of an unpaid central library as a separate use.
How Project Panama worked
Anthropic changed its acquisition strategy as its lawyers grew concerned about training on pirated material. In February 2024, the company hired Tom Turvey, who had led partnerships for Google's book-scanning project, and tasked him with obtaining a comprehensive book collection while limiting the legal and commercial work associated with acquiring it.
Turvey briefly approached publishers about licensing books for AI training, according to the court order. Anthropic then pursued bulk purchases through book distributors and retailers.
The physical workflow was destructive by design:
- Anthropic purchased print books in bulk, including used copies.
- Contractors removed each book's binding and cut apart its pages.
- The pages were scanned and converted into searchable PDF files.
- The physical copies were discarded rather than retained alongside the scans.
- Anthropic kept the resulting digital copies for internal use.
An internal planning document called the effort Project Panama and said Anthropic wanted to "destructively scan all the books in the world." The document also indicated that the company did not want the effort publicly known, according to The Washington Post. A July 27th post on X by Ole Lehmann (@itsolelehmann) later recirculated language from the document.
The available records describe millions of dollars spent on millions of physical books, but they do not provide a more precise total for Project Panama's purchases or scans.
Why the physical scans received different treatment
Alsup held that training language models on copyrighted books was a transformative fair use. He separately found that converting a lawfully purchased print book into an internal digital copy was fair use when Anthropic destroyed the original and did not distribute the digital version.
The physical process preserved a one-for-one relationship between a purchased book and its internal replacement: Anthropic bought a copy, destroyed that copy during scanning and retained the resulting file. The ruling therefore treated the scan as a format change rather than the creation of an additional unpaid copy for distribution.
The pirate-library downloads did not receive that protection. Alsup found that Anthropic had acquired those files without paying and retained them in a general-purpose research library. Building that library was a separate, non-transformative use, even if model training itself could qualify as fair use.
That split explains why Project Panama's purchase-and-scan workflow survived legal scrutiny while Anthropic faced potential statutory damages over its earlier downloads.
What the $1.5 billion settlement covers
The final approval order covers 482,460 books acquired from LibGen and Pirate Library Mirror. At roughly $3,000 per book before allocation adjustments, the agreement resolves claims arising from Anthropic's past acquisition and copying of the covered works.
Anthropic must destroy the original files downloaded from LibGen and Pirate Library Mirror, as well as copies derived from those repositories, subject to requirements that evidence be preserved.
The settlement is not a blanket resolution of every potential dispute between Anthropic and authors. Its release covers claims arising from acquisition and copying through August 25th, 2025. It does not release claims based on model outputs or Anthropic's future conduct.
Anthropic has maintained that the pirate datasets were not used to train any commercially released model. Deputy general counsel Aparna Sridhar told the Associated Press that the June 2025 ruling established AI training on books as fair use and that the company was preparing to close the matter.
The case ultimately produced two distinct results. Anthropic's founders secured judicial support for training on books and for internally digitizing lawfully purchased copies under the process used for Project Panama. But the company's decision to build and retain a library from pirate repositories generated a $1.5 billion settlement and a requirement to destroy the covered files.