build1 distinct publisher
Ten agent eval protocols say when a run stops. Fewer say whether the result is settled.
A replay that held the agent's actions fixed produced different labels before and after delayed operations resolved, and one late write moved the following run's score.
Publishers:arxiv.org
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap−8
- Incentives34
- Confidence50