Frontier models average 59.5 on Argo-Bench and clear 95 on only 34.8% of its 210 enterprise data tasks. The benchmark grades the bans and refunds an agent files against a hidden simulator, a step outside what text-to-SQL scores measure.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence30
S&P Global Energy says per-dataset Genie agents, bundled by a FastMCP proxy, cut launch time for conversational data products from months to days. Its governance claim rests on Unity Catalog, though part of the estate sits in sources outside Databricks.
Reality
- Evidence30
- Adoption25
- Hype gap+35
- Incentives85
- Confidence55
A dev.to post counts the questions whose correct answer needs a table the caller may not read. On a 42-object demo schema with an ordinary role split, the analyst's share is 38.5% and the CFO's is zero.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+22
- Incentives68
- Confidence42
Name-only BM25 matched nothing for "money we gave back to shoppers". The description channel put the right table second and embeddings put it twenty-second, and reciprocal rank fusion let the blind channel cost the answer nothing.
Reality
- Evidence64
- Adoption15
- Hype gap−10
- Incentives60
- Confidence58
A dev.to post argues that logging prompts, SQL and latency cannot show how an agent read the question, and proposes eight append-only events running from resolved intent to final answer. It works if your metrics are already governed.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence55
A buyer's checklist published on dev.to puts accuracy measurement and permission enforcement ahead of the live demo, and every question on it comes with a test you can run inside the meeting. Its author sells in the category.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives75
- Confidence45
The generated descriptions were accurate and specific, and indexing them beside the table names dropped BM25's IDF for the term contact to 0.15, about what the same index gives words the tokenizer never strips out.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+12
- Incentives32
- Confidence57
Yang Fei and three co-authors add roles and column-level policies to Spider, BIRD and LiveSQLBench, then score existing systems on a failure class that covers queries returning the right rows to a user barred from the column.
Reality
- Evidence42
- Adoption8
- Hype gap+25
- Incentives65
- Confidence45
The published code treats INTO and FOR UPDATE as the writes they are, and it walks CTE bodies and union arms for table references that the top-level FROM clause hides. The tenancy fix is the part still to verify.
Reality
- Evidence62
- Adoption16
- Hype gap+12
- Incentives65
- Confidence55
Three LoRA stages on Qwen2.5-0.5B pushed reward to 1.0 and accuracy below the untrained base. The fix that recovered 43 points was an execution harness that runs both queries and compares the rows.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+18
- Incentives38
- Confidence55
The text-to-SQL repo went read-only on 29 March 2026 with its last push in February, but the README says nothing and the cloud product still sells. If you pinned it, the code is fine and the perimeter is yours.
Reality
- Evidence74
- Adoption48
- Hype gap−12
- Incentives62
- Confidence66
Thinking Machines and four academics trained one model past the usual agent pipelines on text-to-SQL, but the part worth copying is the audit they ran first, which put annotation errors in 61.1% of the benchmark examples they checked.
Reality
- Evidence57
- Adoption20
- Hype gap+14
- Incentives70
- Confidence46
Jedify's CTO attributed the misses, re-ran part of its own benchmark on open weights, and owned a bad caption. The arithmetic on its ROI multiples still does not close.
Reality
- Evidence52
- Adoption16
- Hype gap+24
- Incentives78
- Confidence55
A dev.to post argues that recurring natural-language analysis confounds the metric with the instrument. The remedy is a frozen query, reviewed as a diff, with the model on either side of it.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+18
- Incentives30
- Confidence38