Build1 distinct publisher3 min readPublished
Programming a furnace is the eye-catching part, but the piece worth copying is a verification module that refuses any figure it cannot resolve to a logged result, which cuts reported fabrication to 4 percent without telling you what the number means.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Take the cross-check literally. A number in the draft manuscript has to resolve to a value present in the execution log of code the system actually ran, and a figure that does not resolve is treated as fabricated [3]. That is a provenance test rather than a truth test, and its reach is set by one precondition: the number must have been produced by code that executed. The-decoder's account puts the residual rate of fabricated key results at 4 percent with the reliability modules on [20], against the 80 to 100 percent fabrication rates earlier analyses documented in comparable systems [17]. Taken at face value that is a 76 to 96 point absolute reduction [23].
The precondition is where the design gets constrained. In materials science the loop ran through a semi-automated high-temperature furnace, and Co-Scientist generated growth recipes tailored to that lab's equipment for a 2D material previously produced mainly by hazardous etching [6]. Twenty-five rounds of human refinement later, the team had layered structures whose properties resemble the target, and definitive confirmation of the atomic structure is still pending [7]. That is a claim no execution log can settle, and the instrument that could confirm it has yet to weigh in.
The thin-film run pushed more of the physical world into the trace. Gemini 3 Deep Think drove the equipment directly, three semiconductor films came out on the first attempt, and recipe development went from days to minutes [8]. Humans still loaded samples and precursor materials by hand, and the fast mode produced smaller, less uniform crystals than carefully optimised recipes would [9], which is the sort of detail that decides how much weight the phrase "closed loop" is carrying. Lead author Samuel Schmidgall writes that whether those recipes transfer to other labs remains open [10].
Agent_H is a useful limit case for the module. Its advantage on health benchmarks came from automated evaluators [16], and after correcting for overly long responses it beat six frontier models including GPT-5 and Claude Opus 5 [14]. Those are exactly the kind of numbers a provenance check waves through. Three board-certified physicians scoring nine categories found a statistically significant advantage over baseline Gemini 3.1 Pro in one of them, a lower risk of potentially harmful responses [15], and the automated evaluators correlated only weakly with the physicians' judgments [16]. For the benchmark result to transfer to a clinic, the automated grader would have to track clinician judgment, and here the two diverged.
The review study is worth sizing. Thirty domain experts filed 450 independent reviews across 150 autonomously generated papers [19], which works out to three reviews per paper [21] and fifteen reviews per expert [22]. Three reviewers per paper is a serious audit by publishing standards and a thin net for subtle fabrication.
The biology leg is the honest one about scope. Predictions matched unpublished lab results on three of four shape features [11], and the researchers state the system reasons between known conditions and cannot predict behaviour in entirely new systems [12].
If I were adding one thing to a report-writing agent this quarter, it would be the log cross-check with the fabrication and plagiarism penalty next to it [18]. The cost is stored logs addressable per claim and a writer component that may not restate a figure from memory. What you get back is certified provenance, which is all it claims to be.
Ranked by verification strength, evidence, and original report placement.
Verification modules cross-check numerical claims in the text against the execution logs of the generated code in order to cut down on fabricated results.
Co-Scientist addresses fabrication two ways: the system is penalized for fabricated or plagiarized content, and a separate verification module cross-checks every numerical claim against the actual results of the executed code.
With the reliability modules active, Co-Scientist fabricated key results in 4 percent of cases, per the-decoder.com's report of the study.
Google DeepMind expanded its multi-agent system Co-Scientist from a hypothesis generator into a lab-integrated research partner built on current Gemini models, which now plans experiments, writes code and controls lab equipment.
The new element is a closed-loop research workflow: the system derives hypotheses from a research question, creates experimental plans, programs or machine-readable lab protocols, analyzes results, and generates scientific manuscripts.
Google first introduced Co-Scientist in February 2025, then based on Gemini 2.0, with shortcomings in fact-checking and literature review.
Distinct publishers with included, body-backed reporting in this cluster.
Follow any of these and your For You feed starts watching them — no settings page required.
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
invest
Korea cuts one of four sovereign AI teams, and usability did the cutting1 distinct publisher
invest
DeepMind now adds two researchers for every one it loses. In 2023 it was twelve.1 distinct publisher
build
OpenAI Wrote The Hazard Notice Itself, And English Employment Law Knows What To Do With One1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One newsroom relaying one lab's homework
The specificity is real - named lead author, named models per experiment, an ablation rather than a bare win rate - and The Decoder volunteers the weak spots instead of burying them. What is missing is any second pair of eyes: no outside lab, no second outlet, no independent read of the study, and the flagship 2D material still lacks structural confirmation. Two numbers that carry the most weight, the definition of a fabricated key result and the identity of the 90 percent comparison system, are unspecified.
Confined to the lab that built it
Every disclosed use is in-house: one furnace in a collaborating materials lab, one E. coli imaging setup checked against unpublished results, one medical architecture graded by three physicians. Nobody outside has run it, nothing is described as purchasable or available, and the lead author says it is an open question whether the recipes survive a move to another lab's equipment. That last admission is the ceiling on this number.
The 4 percent outruns its own definition
Positive, but modestly so, because the reporting does much of the deflating itself. The overstatement lives in the numbers travelling without their footnotes: 4 percent of what, judged by whom, against which comparison system. Underneath it, a benchmark sweep over GPT-5 and Claude Opus 5 collapses to one significant category in front of three doctors, the biology result only interpolates between conditions it has already seen, and Schmidgall concedes the system writes plausible methods sections that do not describe the code it ran. 'Plans experiments and runs lab equipment' is true; a human still loads the precursors.
The graded and the grader share a payroll
DeepMind built the system, designed the reliability modules, ran the double-blind study that scored them, chose the comparison system and defined the safety filter that rejects 98.7 percent of bad directions. The published caveats are creditable and cut against interest, which is why this is not higher. The clock matters too: OpenAI is reported to be showing an intern-level research agent this fall, so a strong fabrication number lands in the middle of a race for the automated-science narrative.
Firm on what was said, thin on what is true
We can be fairly sure what the study reported, because the reporting is precise and internally consistent. We can be much less sure the reported effects mean what the percentages suggest, since a single outlet, a single vendor study and undefined key terms all stack in the same direction. Our read would move quickly on two things: a second account of the study, or any lab outside the collaboration running these recipes.