Build1 distinct publisher3 min readPublished
The three track leaders have no manifest anyone can download. The curl Bench'd publishes for independent verification points at a host that does not resolve. The badge behind the board bills from $299 a month.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
A leaderboard row is arithmetic over dimension scores. That is the only reason the format carries any authority: anyone can redo the sum. When the sum does not reproduce, the row has stopped being a claim about the vendor and become a claim about the harness. On the recomputation published in the dev.to piece, Bench'd's 80.0 sits over dimensions averaging 26.7, roughly a factor of three out [3][20], and its 60.0 sits over dimensions that are all zero [3]. The 100.0 appears beside a reliability score of 4.0, which the author reads as the fingerprint of an adapter echoing expected answers back to the scorer [c3b]. A design that averages a content score with a reliability score, rather than letting reliability veto the total, will rank that echo first every time.
The next thing you reach for is the manifest, and none of the three leaders has one anyone can download [6]. The receipt is next: the methodology page says "You don't need to trust us. Every receipt can be independently verified," and hands you a curl against benchd.dev [7]. That host answers NXDOMAIN from 1.1.1.1, from 8.8.8.8, and from the author's own resolver, while benchd.ai resolves normally [8]. The independent-verification path terminates at a hostname nobody stood up, and establishing that took four seconds [c8b].
The third fallback is signatures. The trust page defines authenticity as a signature from a key published on that page, and the key that signed Bench'd's shipped manifests matches neither published fingerprint [9].
The most useful figure on the board is the one Bench'd put there as a sanity check. It runs GPT-4o-mini with no memory attached as a control, and on 2026-08-29 that control finished third in the Conversational Memory track [14]. Two ranked memory systems are therefore indistinguishable from having no memory, on the axis the track is named after.
The author's own reporting is the format worth copying: RE-call at 69.0 LongMemEval and 71.6 LoCoMo, unmodified scoring path, $6.64 of spend, on 2026-08-23 [17], published next to a no-memory control of 57.6 [18]. Subtract and you get 11.4 points [19]. The rank does not transfer to your workload; the delta might, if your traffic resembles LongMemEval's, and that is at least a question you can put to it.
The upstream defect is more mundane. One project's published 96.6% recall came back as 87.0% when rescored with the benchmark's own scorer, because it reported recall_any where the spec says recall_all [15]. That is 9.6 points from one word in a field name [21]. The ground truth underneath is not clean either: per the audit by @dial481 that the piece cites, LoCoMo carries roughly 99 wrong or misattributed answers, and a judge accepted about 63% of deliberately wrong responses [16].
So the test for a row in any of these tables is a manifest you can fetch, dimensions that reproduce the total, and a control in the same table. Bench'd's top three clear neither of the first two, and where its control does appear it beats two of the systems ranked above it [14].
Ranked by verification strength, evidence, and original report placement.
Bench'd (benchd.ai) calls itself the neutral benchmark authority for AI memory and sells vendors a verification badge priced from $299 to $3,999.99 a month.
RE-call scored 69.0 LongMemEval and 71.6 LoCoMo through Bench'd's harness on the unmodified scoring path on 2026-08-23, at $6.64 total spend.
Bench'd's no-memory control scores 57.6 on the same scoring path.
The author ran their own memory system through Bench'd's harness, and read every number on Bench'd's site and repository on 2026-08-29.
Bench'd's methodology page states that rows marked Community-Verified are "independently run" by Bench'd, and all three track leaders are marked Community-Verified.
None of the three Bench'd track leaders has a manifest that anyone can download.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 29, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Disproving one pointer in Lemmalog retracts every conclusion that rested on it1 distinct publisher
build
Memora grades memory agents on what they should have forgotten, and six of them fail1 distinct publisher
build
AgentCL: if the task stream is not controlled, agent memory gains prove nothing1 distinct publisher
build
A GPU SQL Engine Lost to One CPU Thread Because a Dispatcher Constant Was 128x Too Small1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One auditor, cheap checks, no rebuttal
The documentary half of this is strong and unusually cheap to redo: benchd.dev returning NXDOMAIN against three resolvers, Bench'd's trust page promising it takes no vendor money next to a pricing page billing vendors, a harness repo with no commit since 6 June and five vendor issues nobody has answered. The load shifts when you reach the parts that need artifacts — the recomputed dimensions averaging 26.7, the signing-key fingerprint mismatch, the assertion that Letta and gbrain never submitted those rows. Those are stated, not shown, and the rows themselves have no manifest to check them against. Bench'd is unquoted throughout, and the author sells a competing memory system.
A board vendors argue with and nobody minds
There is real, if thin, traffic around Bench'd: two vendors have filed five issues on the harness, the author installed it and got a full scored run out of it at $6.64, and three vendor names sit on the front of the board. What is absent is anyone on the other side of the transaction — no badge customers, no traffic or citation counts, no buyer who says a Bench'd rank shaped a decision. Sold-but-unattended is the pattern the evidence actually supports: pricing tiers up to $3,999.99 a month, a repository untouched since June, and a dispute filed on 23 August still unanswered six days later.
Tight arithmetic, field-sized conclusion
On Bench'd specifically the reporting mostly earns its temperature — a verification promise pointing at a host that has never existed is exactly as bad as it sounds, and the author flags his own conflicts before anyone else can. The overshoot is in the widening: from one leaderboard read on one afternoon to 'every project I checked' and an implied base rate for the field, carried by an unnamed 96.6%-to-87.0% rescoring and a second-hand LoCoMo audit. The word 'misrepresented', applied to two named companies on inference alone, is also doing more work than the evidence behind it.
Everyone in frame is selling
Both sides of this have a till. Bench'd ranks vendors and invoices them for the badge that advertises the ranking, from $299 to $3,999.99 a month, while its trust page says it takes no vendor payment — the author's line that an organisation which ranks vendors cannot also invoice them is the cleanest sentence in the piece. And the auditor is a competitor: RE-call's builder publishes RE-call's scores, discloses that he set the abstention threshold to zero and compares only inside one track, then closes by asking readers whether he should have matched rivals' softest metrics instead. Disclosed self-interest is still self-interest.
Reproducible in principle, unreplicated in fact
Halfway, and for a specific reason: the checks are described precisely enough that a second party could settle most of them this afternoon, and none has. One publisher, one author, one reading date, no response from the organisation being accused and none from the two vendors named in the disputed rows. Where output is pasted in — the resolver failures — confidence is high; where a fingerprint or a dimension table would be needed, it drops to the author's word.