science1 distinct publisher
A 91-task benchmark for agent knowledge work runs every task cold, and says so
Artificial Analysis' AA-Briefcase grades deliverables built from four multi-week projects, but each task starts with no memory of the model's own earlier submissions. Continuity stays unmeasured.
Publishers:artificialanalysis.ai
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+12
- Incentives62
- Confidence55