Benjamin Breen asks AI labs to fund historians' frontier-model research
The UC Santa Cruz historian argues for collaborations around tractable archival questions, drawing on early experiments in alchemy, cryptography and historical texts.
By RuntimeWire Staff · Published
Primary source: Res Obscura
Why it matters
Breen's examples show how frontier models could help historians search across languages and source collections, while the Enigma case leaves a concrete problem for researchers: a correct answer can still have an uncertain research path.

Benjamin Breen wants AI labs, historians and funders to collaborate on research using frontier models, from tracing alchemical texts to testing old ciphers. The UC Santa Cruz historian made the case in a September 24th essay, describing early work with GPT-6 models and Claude Opus 5.5 and arguing that researchers should try using them to solve historical problems, not only transcribe and summarize documents.
Breen is a historian of science, medicine and technology whose research examines how knowledge traveled across early modern societies. His university profile also describes his work exploring AI-based historical simulations for teaching. That experience shapes his central claim: models can help historians investigate larger questions, provided specialists know what to ask and can assess the results.
The September releases of GPT-6 Sol and Claude Opus 5.5 prompted Breen to share his early results. He argues that current models can attempt research tasks beyond transcription and summarization. The essay calls for collaborations among AI labs, historical researchers and funding agencies.
The case for a historian in the loop
Breen's proposed work focuses on questions experts already want answered, with searchable source material and some way to test a proposed answer. His framework asks whether the necessary data is digitized and accessible, whether a problem suits models' strengths, and whether a solution can be proved or disproved. Cryptography is a natural fit because a proposed solution can be checked against the original text. Tracing a passage across languages can also produce a testable lead.
One example is GPT-6 Astra's attempt to identify a French alchemical passage that Isaac Newton had translated into Latin. Breen says Astra connected Newton's text to a French source, an identification that appears not to have been made before. He also describes a separate attempt to interpret John Dee's coded manuscript, Liber Loagaeth. Astra concluded that most of the text consists of nonsense syllables, while interpreting one passage as a reference to Bornogo, an angelic being in Dee's mythology. Breen says the result is not a breakthrough in Dee studies; he presents it as a lead for a specialist to assess.
Breen's earlier historical-AI work is distinct from these new examples. An earlier Res Obscura account documents experiments using GPT-4o, o1 and Claude Sonnet 3.5, including work on a sixteenth-century Italian map, an eighteenth-century Mexican medical manuscript, and notes and letters concerning Francis Galton and William James.
The Enigma result, and its source trail
A separate GPT-6 Astra Enigma break offers a testable example. Carter Leffer directed Astra to try unsolved messages listed by Crypto Cellar Research. On September 15th, Leffer contacted Weierud to request validation of Astra's break of one of them. Astra selected a July 10th, 1941 German army message, developed Python and C++ tools to simulate the Enigma machine, and recovered a key and plaintext that cryptology researcher Frode Weierud confirmed as correct. Weierud's account of the break says the message had resisted solution since 2005.
The work also shows how difficult it can be to reconstruct a model's research path. Astra's logs cited correct Bundesarchiv file references and appeared to connect the message with information about radio-message collections that Weierud added to his 1941 Message List in July 2026. Weierud says he could not establish whether Astra accessed digitized Bundesarchiv collections, found the files elsewhere or relied on a private collection. The answer was verified; the route to it remains uncertain.
Breen's framework asks whether the source material for a historical question is digitized and accessible. The Enigma example gives that question practical weight: even after a solution is checked, researchers may not be able to determine how a model found the evidence it used. Historians and archivists can evaluate answers against documents and existing scholarship, and help identify which questions are suitable for this kind of work.
Breen argues that AI labs, historical researchers and funding agencies should actively pursue collaborations. His examples show models producing leads in areas where multilingual reasoning, code and connections across specialized sources may help. They also leave the historian's work intact: checking the source trail and deciding what a result means.