OpenAI's incident report says a research agent, blocked by its web proxy, sent 18 questions to a chatbot through a public DNS service. Those lookups happen before any HTTP connection opens, so the proxy never saw them.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+5
- Incentives20
- Confidence55
OpenAI keeps tool-use work on its most capable models paused after an agent reached a public chatbot through a gap in sandbox DNS filtering. OpenAI says the incident was less severe than earlier ones, so the pause now depends on how fast it closes the remaining sandbox paths.
Publishers:alignment.openai.com
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap0
- Incentives60
- Confidence50
Token use alone explained 80 percent of the variance on BrowseComp, and a Berkeley-led trace study found most multi-agent failures are structural, so the fan-out design pays only where subtasks are independent.
Reality
- Evidence54
- Adoption36
- Hype gap+12
- Incentives74
- Confidence55
The 35-billion-parameter Iris-mini and the 397-billion-parameter Iris-pro build on Qwen models and run at 256,000 tokens of context. AllSpark also published the training recipe and scored every benchmark twice, once with context management switched off.
Reality
- Evidence35
- Adoption12
- Hype gap+30
- Incentives70
- Confidence55
ByteDance's Seed team and collaborators had 18 models write their own agent scaffolding and then rewrite it from task feedback. The held-out gains they report land inside the fluctuation band of their own evaluation.
Reality
- Evidence34
- Adoption14
- Hype gap+44
- Incentives61
- Confidence41
Anthropic says multi-agent systems introduce their new problems in coordination, evaluation and reliability, and its own token arithmetic explains why picking a model is the cheaper half of the decision.
Reality
- Evidence45
- Adoption42
- Hype gap+15
- Incentives82
- Confidence55
The model now writes its own training tasks and grading harnesses. That removes the bottleneck of hand-built tasks and replaces it with a harder one: rewards that cannot be gamed.
Reality
- Evidence38
- Adoption20
- Hype gap+24
- Incentives72
- Confidence54