Red Hat's post puts generative AI's productivity gain on the hardest enterprise codebases at 0-10 percent, then offers a pipeline that grounds the model in migration rules a customer's own team has to write and curate.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+35
- Incentives85
- Confidence55
CEO Itamar Friedman says the ceiling exists to force someone to answer which automation is worth the money. The bill growing faster is the infrastructure Qodo runs for customers, up roughly 5x year over year.
Reality
- Evidence34
- Adoption28
- Hype gap+22
- Incentives76
- Confidence43
An agent drafts the CSV import in an afternoon; the week of making it correct is what remains billable. Three large surveys say what that week is worth. A controlled study says developers get their own hours wrong.
Reality
- Evidence55
- Adoption45
- Hype gap+25
- Incentives70
- Confidence50
Daily agent use among engineers climbed to roughly 80% in a year while developer trust in the output fell to 29%. The case that the gap is costing delivery time rests on a randomized trial of 16 people.
Reality
- Evidence45
- Adoption74
- Hype gap+28
- Incentives72
- Confidence50
A dev.to post argues that the approve and block labels AI review tools print are uncalibrated text, and proposes fitting sigmoid(alpha + beta * verdict) to a team's own review history so triage gets a threshold it can defend.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence45
A field study of 4,867 developers recorded 26.08% more tasks completed with GitHub Copilot, while METR timed 16 experienced developers taking 19% longer in repositories they already knew. The two figures count different things.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives55
- Confidence40
A principal engineer built the same EV-charging invoice service twice and timed a coding agent through nine cumulative features on each. The two setups differed in more than layering, and he says so.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+12
- Incentives25
- Confidence50
Anthropic's analysis of roughly 400,000 sessions puts about 80 percent of the execution decisions with the agent. The benefit it reports is longer unattended runs between check-ins, counted in actions per turn.
Reality
- Evidence35
- Adoption45
- Hype gap+30
- Incentives70
- Confidence45
The step in front of the person got faster. The steps behind it did not, so an afternoon of work aged in review for months. A survey of 4,500 executives in nineteen countries describes the same condition at scale.
Reality
- Evidence58
- Adoption38
- Hype gap+12
- Incentives68
- Confidence55
GitClear's Maintainability Gap report measures output and code health on the same 623 million changes. The two move in opposite directions, and the denominators cover different teams.
Reality
- Evidence48
- Adoption45
- Hype gap+18
- Incentives65
- Confidence55
A consultant's six-month log across forty shipped Python automations reports build time 38 percent below baseline and client revisions halved, with an explicit warning that React work sits outside the sample.
Reality
- Evidence26
- Adoption21
- Hype gap+33
- Incentives55
- Confidence52
Cursor's Projects beta gives each body of work a coordinator agent on its own cloud machine. Its internal deployment shows where the review load lands, and its productivity figures arrive without a baseline.
Reality
- Evidence44
- Adoption36
- Hype gap+33
- Incentives78
- Confidence57
Anthropic's usage dashboard reports each team member's email beside their monthly lines of code accepted and exports the list as a CSV, which turns a seat-utilisation tool into something a performance review can reach for.
Reality
- Evidence64
- Adoption
- Insufficient
- Hype gap−12
- Incentives76
- Confidence58
Sean Goedecke argues that visible anger gets you excluded from the conversations where the interesting work is handed out, and the awkward part is that his own essay rates those engineers above average at shipping, so no dashboard catches the loss.
Publishers:seangoedecke.com
Reality
- Evidence26
- Adoption
- Insufficient
- Hype gap+24
- Incentives28
- Confidence61
A 90-day instrumented study reports 28% more pull requests and 13 points of test coverage, but the gain lands only after training and workflow embedding, which puts the money on enablement rather than seats.
Reality
- Evidence45
- Adoption38
- Hype gap+12
- Incentives
- Insufficient
- Confidence40
Andrew Swerdlow says Roblox spent about six months moving from autocomplete to agents, and that the binding constraint is trust and review infrastructure rather than model choice.
Reality
- Evidence32
- Adoption30
- Hype gap+10
- Incentives55
- Confidence42
A theoretical paper assumes error-free, near-free language models and still finds that in two of three deployment scenarios, researchers respond by doing more work with less care.
Reality
- Evidence42
- Adoption28
- Hype gap+18
- Incentives46
- Confidence38
One iOS team's own measurement puts the value of agent workflows in the mechanical half of code review. In Swift, letting agents generate freely moved the time the other way.
Reality
- Evidence24
- Adoption20
- Hype gap+9
- Incentives55
- Confidence41
The multi-agent pipeline is the part everyone will copy. The measurement, and the categories where acceptance falls to 40.6%, is the part worth reading.
Reality
- Evidence58
- Adoption52
- Hype gap+14
- Incentives66
- Confidence54
The reported gain sits in code generation and the reported costs sit in review and post-merge repair. Computed against one unit of work, most teams do not come out ahead.
Reality
- Evidence24
- Adoption62
- Hype gap+42
- Incentives58
- Confidence31