Ai2 releases an open 8B model for cited science reports

Ai2 says AstaBrief averages 51.1 seconds per report in Asta Fast mode, versus 178.5 seconds for Claude-powered Thinking mode.

By · Published

Primary source: Ai2

Why it matters

AstaBrief pairs public model weights and training artifacts with a local-PDF workflow, giving research institutions a path to run cited-report generation on their own infrastructure. Ai2's results are promising but mostly reflect work completed in 2025, and its published tests do not yet establish how well the model preserves the limits of evidence in broader real-world use.

Ai2 releases an open 8B model for cited science reports — AstaBrief averages 51.1 seconds per report in Ai2's Asta system, against 178.5 seconds for its Claude-powered Thinking mode.

The Allen Institute for AI released AstaBrief 8B on October 2nd, an open model that turns research questions and retrieved scientific papers into cited reports, according to Ai2's announcement. The release extends the open-model mission Paul Allen established at Ai2, the Seattle nonprofit he founded in 2014, into a practical research workflow: Ai2 says AstaBrief now powers Fast mode in its Asta platform, alongside a Claude-powered Thinking mode.

Allen, the Microsoft co-founder and philanthropist, founded Ai2 to develop AI for problems he considered consequential. AstaBrief carries that institutional bet into a specific, difficult task: helping researchers synthesize literature without letting a fluent answer outrun the evidence. The model release is credited to Ai2 as an organization, and its announcement frames the work as part of a broader program to make AI for science open and adaptable.

The speed claim is about the whole workflow

Ai2 reports in its announcement that Fast mode takes an average of 51.1 seconds per report across the Asta pipeline, compared with 178.5 seconds for Thinking mode, or roughly 3.5 times faster. Ai2 also describes the reduction as nearly an order of magnitude relative to proprietary models it tracked, but the headline pipeline figures show a smaller difference and do not establish a tenfold speedup for Asta users.

The design change helps explain the timing. Fast mode takes a question and relevant retrieved excerpts, then generates the report in one pass. Thinking mode uses additional steps to summarize and cluster excerpts before writing the report section by section. AstaBrief therefore handles report generation; literature retrieval happens upstream. The two modes are options within one product, with the slower pipeline retained for more compute-intensive work.

Diagram comparing Asta's Fast and Thinking report-generation workflows after literature retrieval.
Ai2 describes Fast as a one-pass report workflow and Thinking as a multi-step workflow that summarizes and clusters excerpts before writing - AI explanatory diagram, not documentary evidence. RuntimeWire - AI-generated diagram.

Ai2 says in its training account that the model is built on Qwen3-8B and trained with supervised fine-tuning and direct preference optimization, rather than reinforcement learning. The institute's stated reason was operational: it wanted a less expensive, easier-to-debug training recipe, with data quality doing more of the work. The team began with 90,000 research-focused queries filtered from user logs, generated candidate reports with several proprietary models, and retained 47,000 examples for supervised fine-tuning. A separate preference-training stage used about 6,000 report pairs.

That is an open release built partly from outputs generated by closed models. Ai2 says reports used to create supervised examples came from Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini and GPT-4.1. For preference data, the institute compared reports generated by different systems and kept examples where two judge models agreed on the preferred report. The approach makes the training pipeline inspectable, while also showing how open models can be shaped using proprietary systems as teachers.

Diagram of Ai2's separate supervised fine-tuning and preference-training data paths for AstaBrief 8B.
Ai2 says supervised examples used reports from five named proprietary models, while preference examples were retained when two judge models agreed - AI explanatory diagram, not documentary evidence. RuntimeWire - AI-generated diagram.

Open weights meet a narrow evidence test

Ai2 has released the AstaBrief model, the DPO dataset and an example workflow for generating reports from local PDFs. The model card lists an Apache 2.0 license and describes the model as taking retrieved literature excerpts as input. Ai2 says local deployment can help institutions handle research questions involving sensitive or unpublished work without sending report-generation requests to a proprietary model API.

The released evaluation shows progress on a bounded test, not proof that the model reliably handles every scientific field or real-world research question. On the model card's ScholarQA-CS2 test set, consisting of 100 computer-science research questions, AstaBrief scored 87 overall, with 90.5 for citation precision and 78.2 for citation recall. Ai2's own post warns that most training and evaluation work was completed in 2025 and that it has not rerun the full comparison against 2026 frontier models. Its scores should be read as evidence for the training and system choices it tested, not a current ranking against today's proprietary models.

The remaining scientific challenge is larger than attaching citations. A report can cite a relevant paper and still overstate what that study found, generalize beyond its sample, or turn an observation into a recommendation. Ai2 says its main development metrics measured relevance, coverage and citation grounding, and identifies preserving the scope and strength of source claims as an area that needs better evaluation.

Early Asta usage offers a product signal, with limits. Of 374 users who tried Fast mode, Ai2 says 29.1% used it on at least two days; 23% continued with Fast mode without switching back to Thinking mode, while another 18% alternated between modes. Positive feedback rates were 84.2% for Fast and 85.2% for Thinking, though Ai2 cautions that feedback was too sparse for strong conclusions. Those figures describe a small group of users who tried the feature, not Asta's overall adoption or paid demand. Ai2 reports these usage figures in its announcement.

The release fits Ai2's longer-running strategy of putting model weights and research artifacts into public hands. On October 1st, RuntimeWire reported that Ai2 had opened parts of its Olmo training stack; AstaBrief applies the same institutional preference for openness to a narrower layer of the scientific AI workflow. For researchers, the useful bet is control over deployment and the ability to inspect and adapt the model. Whether that control produces trustworthy reports beyond Ai2's tests will depend on evaluations that measure not just whether a sentence has a citation, but whether the cited evidence actually supports its scope.

Reader comments

Conversation for this story loads after sign-in.