Build1 distinct publisher3 min readPublished
Jedify's CTO attributed the misses, re-ran part of its own benchmark on open weights, and owned a bad caption. The arithmetic on its ROI multiples still does not close.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Six SQL assembly errors out of 200 runs is a rounding error everywhere except where these landed. Jedify's own word for that layer is deterministic, and Elimelech says the six are 3 percent of runs and go to the front of the engineering queue [3]. A deterministic component does not have an error rate. It has a bug. The 20 entity-selection failures are a different object: phrasing that legitimately resolves to more than one entity in the graph, the kind of question the company says a human analyst would also have paused over [4]. Those you route to a clarifying prompt. The six you fix.
The comparisons are less settled than the attribution. The release quotes its Token ROI twice, as roughly 4X and 18X in one place and 4.6X and 20X in another [8]. The only multi-agent figure in the release is CHESS at about 339,965 tokens per request [7]; divided by Jedify's 25,036 tokens per SQL-generation call [6], that is 13.6X [3], short of both stated multiples. The schema-injection side behaves the same way. The quoted baseline range of 50,000 to 150,000 tokens per call [7] corresponds to 2.0X through 6.0X against Jedify's figure [4], so the single number sits mid-range in a spread whose ends differ by a factor of three.
Two counting choices pull in opposite directions and neither shows up in the multiple. Jedify's 25,036 includes output tokens while the schema-injection baselines count schema input alone [10], which understates the gap. Caching pulls the other way, by an amount nobody has published, which is why the first draft's caption mattered: it described 25,036 as an effective count after prompt caching, when the number came raw out of API logs with no discount applied [9]. The figure survived review; the explanation of it did not [9]. A cache-parity study is the committed follow-up [11].
The 85 percent routing projection was the softest thing in the report, resting on dbt's finding that model choice matters less once context is scoped properly [12]. It now has a measurement attached to it: two to five accuracy points lost on an open-weight model across the simple and moderate tiers [13], which puts those tiers at 82 to 85 percent against the reported 87 [5]. The complex tier has not been run [13]. That is the honest shape of the result, and it leaves the expensive half of the question open, because complex queries are where an open-weight substitution is most likely to fail and most costly when it does.
What is worth copying here is the procedure rather than the numbers, which remain company-reported [16] against baselines drawn from other people's schemas and question sets rather than a same-warehouse run [7]. Asked to account for the failures, the vendor finished the attribution and ran a new experiment [1][2], and says the next revision will carry the measured number instead of the architectural argument [15]. The remaining assertion is Elimelech's claim that non-interactive grading makes 87 percent a floor rather than a ceiling [14]. That one is testable, and it is the obvious candidate for the next run.
Ranked by verification strength, evidence, and original report placement.
The context graph approach, which pre-encodes business logic instead of feeding raw schema to the model, answered 87 percent correctly while averaging 25,036 tokens per SQL-generation call.
Adi Elimelech, co-founder and chief technology officer of Jedify, answered Lets Data Science in writing about the company's Context Graph benchmark, published the same day.
Asked about the 26 failed runs the report left unattributed, Jedify finished the attribution: 20 were entity-selection misses on genuinely ambiguous questions and 6 were SQL assembly errors in the layer Jedify describes as deterministic.
Elimelech called the six SQL assembly errors "the ones I like least, because they touch the layer we describe as deterministic", said they are 3 percent of runs, and said "they get engineering attention first".
Twenty of the 26 failures trace to entity selection, on questions whose phrasing "legitimately resolve[s] to more than one entity in the graph, where a human analyst would also have paused to ask a clarifying question".
The study ran 100 business questions against a live production data warehouse, each twice, for 200 graded runs across three complexity tiers.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but vendor-run and single-sourced
There is unusually granular primary evidence for a vendor benchmark: 200 graded runs, a completed failure attribution, an on-the-record caption correction, a disclosed build cost, and a fresh open-weight re-run performed on request. But every number originates with the vendor, the comparison baselines come from different schemas and question sets rather than a same-warehouse run, the ROI multiples are stated two inconsistent ways, and the cache-parity and complex-tier results do not yet exist.
One vendor-run environment, no external uptake disclosed
The only concrete deployment evidence is the single live production data warehouse used for the benchmark, encoded as a 35-entity graph in about three days. No customer count, third-party deployment, usage volume or out-of-coverage rate is disclosed, and the vendor explicitly declined to estimate what share of production questions fall outside the graph.
Mildly overstated multiples, unusually candid framing
The marketing arithmetic runs ahead of the data: the cited CHESS figure divided by Jedify's own per-call number gives about 13.6X against stated multiples of 18X and 20X, the schema-injection band implies anywhere from 2.0X to 6.0X, and the 85 percent open-source routing figure is measured only for the easier tiers. The gap is kept modest by the vendor's own disclosures, correcting a wrong caption, conceding an instrumentation gap, declining to invent a coverage figure, and a counting asymmetry that works against its own numbers.
Vendor-authored benchmark promoting its own architecture
The underlying study is published by Jedify about Jedify's context graph product, and the interview subject is a co-founder with a direct commercial interest in the token-cost and accuracy narrative. Same-day publication of report and Q&A, plus ROI multiples stated in marketing terms, sharpen that incentive; the vendor's willingness to publish failure attribution, build cost and a correction partially offsets it, but does not remove the self-interest.
Strong primary sourcing, single publisher
Confidence rests on a direct written Q&A with the responsible executive, quoted at length, with the publisher flagging its own limits. It is capped by there being one publisher, no independent replication, unnamed models in the re-run, no disclosed grading methodology, and two committed follow-up studies still outstanding.
leadership
EY Answers The AI-ROI Question With An Org Chart: One Office, One Budget1 distinct publisher
build
EY turns AI cost control into a standing office, and claims 60% fewer tokens for it1 distinct publisher
build
Duolingo's video call went from 30 cents to a penny, and took a pricing tier with it1 distinct publisher
build
Grok 4.6 lands on Bedrock at $2/$6, turning an xAI decision into a line item2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 26, 2026