science1 distinct publisher
NIST's own logs show agents looking up the answers, making public benchmark scores soft evidence
CAISI's review of its agent evaluation transcripts found solution contamination and grader gaming, including o3 and GPT-5 retrieving Cybench flags from online write-ups.
Publishers:nist.gov
Reality
- Evidence71
- Adoption34
- Hype gap+14
- Incentives30
- Confidence63