Build2 publishers2 min readPublished
Ai2's open AstaBrief 8B writes cited research reports 3.5 times faster than Asta's Claude mode
Ai2 open-sourced AstaBrief 8B, which writes cited research reports in 51.1 seconds against 178.5 for Asta's Claude-powered mode. Labs that cannot send unpublished research questions to a hosted model can now run a cited-report generator on their own servers.
The Engineer · Build desk
What happened
- AstaBrief already runs as Fast mode in Asta's Generate a report feature, and Ai2 published the training data along with the weights.
- Ai2 started from Qwen3-8B and put most of its effort into post-training data, evaluation and the report-generation scaffolding.
- Training drew on tens of thousands of real research queries, filtered with a focus on citations and paired with preference data.
- Ai2 also released an example workflow that researchers can adapt to build reports from their own PDFs on local machines.
- Most training and evaluation finished in 2025, and Ai2 says it has not rerun the full evaluation against today's frontier models.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost A lab that wants to retrain or extend AstaBrief can use the published data with SFT and DPO and skip the judge-in-the-loop RL setup that Ai2 says can be unstable and expensive.
- decision Teams choosing between AstaBrief and a hosted model need their own citation-grounding comparison against current models, because Ai2's parity test stopped at the 2025 frontier.
- constraint The 51.1-second figure belongs to Asta's pipeline, so a lab running the PDF workflow on its own GPUs changes both retrieval and hardware and has to time the model again.
Ai2 gives two speed figures, and they measure different things. The first, in the post's words, is "nearly an order-of-magnitude reduction in report generation time compared to the proprietary models we tracked" [4]. The second covers the whole Asta pipeline. There, Fast mode averages 51.1 seconds per report against 178.5 for the Claude-powered Thinking mode, so Fast is about 3.5 times faster [3]. Each Fast report arrives 127.4 seconds sooner [1].
Some of the gap between roughly 10x and 3.5x is presumably time spent outside the model, in pipeline stages that a smaller model does not shorten. Some may come from the comparison itself. The larger figure is measured against the proprietary models Ai2 tracked [4], and the smaller one against a mode Asta ships today [2].
The design change I would credit first is in the scaffolding. Ai2 rebuilt report generation to write the full report in one pass instead of section by section [16]. A sectioned pipeline makes several model calls per report, each with its own prompt. One pass makes a single long call over the research question and the retrieved excerpts [1]. The 3.5x figure compares two shipped modes. It does not show how much of the gain came from the scaffolding and how much from the smaller model [3].
On training, Ai2 considered reinforcement learning. Its DR Tulu work showed RL can improve long-form reports from open-weights models, especially with judge models in the training loop [7]. Ai2 chose supervised fine-tuning and DPO instead [7], and I think it chose correctly. "RL-based training can be unstable and expensive," Ai2 wrote [8].
The quality bar covered answer quality, relevance, structure and citation grounding [14]. Grounding matters in Asta because users come back to generated reports later and treat them as working research artifacts [15]. Ai2 framed the project as a test of whether a small open model could match the proprietary models it was using on report quality while cutting generation time and serving costs [13]. Ai2 is careful about what the test shows. Its results, Ai2 wrote, "are best read as evidence about the particular training and system design choices we tested" [10].
What to watch
- A rerun of Ai2's full evaluation against current frontier models would show whether quality parity holds beyond the 2025 comparison set.
- Independent timings of AstaBrief 8B on self-hosted GPUs with the PDF workflow would show how much of the 3.5x gain survives outside Asta's pipeline.
- An ablation separating the one-pass scaffolding from the model swap would show which change bought the speed.