Microsoft Research's Webwright lifted GPT-5.4 from 33.5% to 60.1% on 200 long web tasks by having it write browsing code from a terminal. Because the baseline steered by screen coordinates, the result compares writing code with clicking pixels.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence45
Ramana Kumar used AI to exploit a different bug in each of two Lean kernels, passing off a false Collatz disproof as machine-checked in July. AI labs rely on Lean to vouch for their maths results, so those claims hold only as well as the checker does.
Reality
- Evidence50
- Adoption25
- Hype gap+15
- Incentives55
- Confidence55
IBM and Oxford Economics found that 60% of 8,800 full-time employees worried AI was eroding their skills. The closest controlled test, a 52-developer Anthropic experiment, supports building practice into AI-assisted work and is too small to blame the tools for skill loss.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap0
- Incentives30
- Confidence40
Omni Calculator found 54% of managers run at least 70% of their writing past AI, against 30% of experienced individual contributors. The studies tying that habit to weaker thinking did not study working managers. The case for retraining leaders is still an inference.
Reality
- Evidence45
- Adoption55
- Hype gap+35
- Incentives
- Insufficient
- Confidence40
Microsoft Research's machine-learning pipeline turns solar-wind forecasts into risk estimates for 66,935 U.S. substations, 30 to 60 minutes ahead. Its scores come from forecasts of fast magnetic-field change, so acting on a warning at a particular substation still needs a field test.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+15
- Incentives55
- Confidence40
A Microsoft Research study of mobile manipulation reports slower mapping, later obstacle detection and halved manipulation accuracy on small onboard GPUs. The team has shipped Kubernetes tooling to move that inference off the robot.
Reality
- Evidence52
- Adoption22
- Hype gap+24
- Incentives78
- Confidence46
The Experimentation Platform group compared every pair of A/B tests running on the same day in four products, and argues the interference risk is small enough that isolating tests costs more statistical power than it protects.
Reality
- Evidence58
- Adoption55
- Hype gap+12
- Incentives60
- Confidence52
The sandwich conjecture that Jeong Han Kim and Van Ha Vu posed in 2004 was proved in 2025, so a property established for a random binomial graph now carries to the regular graph containing it, once the graph is large enough.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+12
- Incentives32
- Confidence54
Wharton's Scale of Agency Decay runs experimenting, integrating, relying, depending. The account details the first two, and its usable marker is whether people still overrule a system that contradicts them.
Reality
- Evidence55
- Adoption60
- Hype gap+25
- Incentives55
- Confidence50
Microsoft reports Qwen3.5-9B going from 41.8% to 56.4% on SWE-bench Verified using 6K examples. The load-bearing change is who owns the agent loop, and what the service boundary costs.
Reality
- Evidence32
- Adoption18
- Hype gap+28
- Incentives66
- Confidence44
Microsoft Research says a 4-billion-parameter model trained with SocialRL out-scored GPT-4.1, GPT-5.1 and GPT-5.2 across six bargaining domains. The winning margin is thinner than the method's effect.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+46
- Incentives68
- Confidence33
Microsoft's learned exchange-correlation functional is now in one production code with four more integrations underway. The accuracy claim is the easy part; workflow policy is not.
Reality
- Evidence38
- Adoption27
- Hype gap+28
- Incentives82
- Confidence52