Claude agents designed 1,440 protein binders and 354 bound. The interesting part is that the prompts, provenance and every measurement went out with the number.
Reality
- Evidence60
- Adoption30
- Hype gap+12
- Incentives68
- Confidence62
The framework analyses a paper and its codebase and exposes the workflows as tools a chat client can call, with case studies on AlphaGenome, Scanpy and TISSUE that the authors say reproduce the original results.
Reality
- Evidence60
- Adoption12
- Hype gap+35
- Incentives55
- Confidence58
Essam Heggy and co-authors argue in Nature Sustainability that uncertain drought supply, more than shortage, drives disputes over Africa's 63 shared basins. They want reproducible hydrological models to give negotiators common numbers, a remedy the commentary argues for but does not test.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+30
- Incentives40
- Confidence40
Anthropic's new in-house molecular biology lab says a 21.5 hour agent run turned a 1.9 billion cluster database into 19 reports for humans to read, and one of them described a repeat array nobody had recorded.
Reality
- Evidence46
- Adoption22
- Hype gap+30
- Incentives80
- Confidence60
A cross-sectional reversal baseline on 15-minute Bybit perpetual bars won 67 of 76 out-of-sample windows before costs and none of them after. Charging the trading it requires puts breakeven near 2.5 bps per side.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap−20
- Incentives20
- Confidence45
The zero-knowledge proof in the March 2026 white paper turned out to certify false statements too, and Scientific American reports the algorithm Google was protecting stayed secret for only a few days.
Reality
- Evidence55
- Adoption30
- Hype gap+30
- Incentives50
- Confidence58
A LessWrong post argues that frontier labs publish alignment results with no code and thin methodology. Nobody outside the lab can tell which choices moved the number. It wants a dedicated effort to reproduce them.
Reality
- Evidence33
- Adoption
- Insufficient
- Hype gap+14
- Incentives45
- Confidence46
A Bloomberg developer's agent run recovered an 82-letter German message from 1941 in about ten hours. What makes it checkable is an acceptance procedure that enumerated every doubtful letter before the search started.
Reality
- Evidence62
- Adoption22
- Hype gap+26
- Incentives52
- Confidence60
The offset in a J-lens readout mostly tracks how often a token appears, and scaling it by variance, after plain subtraction failed, lifted hidden-word elicitation to 0.805 from 0.665 on Gemma-2-9B-it. The paired test over 20 words gives p of about 0.19.
Reality
- Evidence46
- Adoption12
- Hype gap+10
- Incentives40
- Confidence56
Jerald Ault and Jiangang Luo spent about three years rebuilding a 1960s federal tagging experiment and put Atlantic menhaden natural mortality at 0.50 a year, well below the 1.17 that entered the SEDAR 69 benchmark assessment.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap−10
- Incentives50
- Confidence55
Paper2Agent tests every function it extracts against the paper's own published output before shipping it. Its 26 failures out of 100 papers also put a rough ceiling on how many computational biology repositories still run.
Reality
- Evidence58
- Adoption22
- Hype gap+8
- Incentives65
- Confidence55
Yucheng Du and Xiyang Hu report a single direction that separates answerable from unanswerable math and code prompts at 0.939 mean AUC across 11 models, and its mean cosine with the canonical safety-refusal direction is 0.087.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+12
- Incentives60
- Confidence50
A molecular optimisation experiment cleared significance on three seeds per policy and produced readable weight matrices. At ten seeds most learned policies trailed the random baseline. The first warning was a ranking that changed with the host.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap−15
- Incentives20
- Confidence55
Chen and colleagues report that a single text-prompted model scored highest overall across 16 pathology test sets, measured against baselines their code list names. The abstract publishes the counts; the margins are left out.
Reality
- Evidence45
- Adoption15
- Hype gap+30
- Incentives55
- Confidence55
Its authors report 30,000 candidate structures and 38 leads in 48 hours from models small enough to run locally, though the comparison with frontier systems arrives in the abstract without a named baseline.
Reality
- Evidence46
- Adoption14
- Hype gap+34
- Incentives62
- Confidence54
A team adapting the implicit association test to reasoning traces found four of five models working harder on association-incompatible prompts, which puts a measurable bias signal in the process rather than only in the answer.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+14
- Incentives45
- Confidence55
A catalytically dead Csm enzyme posted 78 to 90 percent apparent knockdown by RT-qPCR. That result points at guide RNA surviving extraction rather than at any cutting. The primer design that produced it is field convention.
Reality
- Evidence78
- Adoption30
- Hype gap+8
- Incentives35
- Confidence65
The wrapper hashes code and output, sends only a digest out to be signed with RSA-PSS, and hands the reviewer a PDF. The cost of adopting it is a command prefix. The cost of believing it is trusting whoever holds the signing key.
Reality
- Evidence60
- Adoption10
- Hype gap+14
- Incentives55
- Confidence45
Stanford's Le Cong says the fully autonomous lab is a bad idea and that humans should keep the mission. The unsettled question is how AI involvement gets reported to regulators.
Reality
- Evidence24
- Adoption
- Insufficient
- Hype gap+18
- Incentives68
- Confidence36
A Chiba University-led review sets a harmonized nomenclature for selenosugars plus a minimum analytical bar. It is the dull prerequisite for pooling twenty years of urinary selenium data.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+24
- Incentives66
- Confidence57