Build1 distinct publisher3 min readUpdated
A vendor-commissioned test capped every engine at 512 MB. Neo4j Community cleared 10 concurrent clients and returned four of 744 operations at 40, which no latency table would have shown.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
A graph of 2,665 nodes and 103,836 relationships works out at roughly 39 relationships per node [3][19], and nothing about that troubles a 512 MB ceiling [5]. So the thing that ran out at 40 concurrent clients was not the data. Neo4j Community had already cleared the 1-client and 10-client runs with what the tester calls decent latencies [7]. The only variable that moved was the number of clients holding sessions open. The cap itself was picked as one tight enough to be realistic for a free tier [22], which puts the collapse inside the size class a free tier actually hands you rather than in a contrived corner.
Four operations came back out of 744 [17], a 99.5 percent error rate [16] against an attempt rate of about 50 operations per second [18]. The tester's account is that Neo4j's JVM ran out of headroom with 40 threads each opening its own bolt session per operation, and he says plainly that this is a guess rather than a confirmed root cause [9]. Both halves deserve to be taken at face value. The failure is measured; the cause is not.
Memgraph is the control that makes the result legible. It speaks the same Cypher over the same bolt protocol through the same official driver [13], and it has no cliff at all. Degradation and collapse are indistinguishable in a table of medians taken at a single concurrency level, because both look like a bigger number right up until there is no number.
The re-capping decision is the part worth copying. Lifting the comparators to the memory the sponsor's own product was provisioned with, instead of the smaller figure printed in the brief, made the test more generous to them, and Neo4j failed at the larger number anyway. The tester also states that he did not loosen the cap afterwards and did not drop the 40-client run from the results [10], which is the cheapest way to make a number like that go away.
The commission still sits in the frame. CognoDB's worst result anywhere was its own group-by-average aggregation, over three seconds at the median against about a quarter of a second for Neo4j and Memgraph running locally [14], a gap of roughly twelvefold [20] that the writeup attributes in part to the network hop to a cloud instance [15]. Publishing that helps. It also marks where the same-hardware frame thins out, since the engine with no errors was answering over a wire while the ones under the cap were local. The accounting needs one asterisk too: the writeup names five platforms in total [1][2] and then refers to 'all eight platforms tested' [21].
The number worth demanding from a graph vendor is not a median at one client count. It is the client count at which the error rate stops being zero, and the run that shows it.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The writeup refers to CognoDB's aggregation as the slowest result 'across all eight platforms tested'.
The author was asked to benchmark CognoDB Cloud, a managed graph database, against four other graph platforms on the same hardware, the same dataset and the same queries.
The four comparators were Neo4j Community, Memgraph Community, ArangoDB Community and FalkorDB.
The dataset was MovieLens 100K: 2,665 nodes and 103,836 relationships.
Every platform ran the same six measurements: load throughput, 1/2/3-hop traversal latency, indexed lookups, an aggregation query, and a mixed read/write sweep across 1, 10 and 40 concurrent clients.
Every platform was given the same resource cap: 0.5 vCPU, 1 GB disk, 512 MB RAM.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed single-source methodology, no releasable data
The cluster rests on one self-published account by the tester. Methodology is unusually explicit for the format — named comparators, a public dataset with node and edge counts, six enumerated measurements, one shared resource cap, and disclosure of two harness corrections — and the author separates measurement from interpretation by labelling both causal hypotheses as guesses. But there are no raw latency tables, no harness code, no repeat-run variance, no independent replication, and one internal accounting inconsistency about how many platforms were tested. Comparators also ran locally while the sponsor's product ran as a network-attached managed service, an asymmetry acknowledged only in passing.
One commissioned test run; no third-party uptake
The only adoption-adjacent facts supplied are the test itself and incidental deployment details it surfaced: a single five-platform run, one live CognoDB free-tier instance, and a FalkorDB Docker default others had already flagged upstream. There is no usage disclosure, customer count, deployment story or downstream reaction from any of the five vendors, and no sign that the results changed anyone's platform choice.
Mildly overstated by framing, restrained in substance
The substance is more disciplined than the packaging. The author retains the adverse run, publishes the sponsor's worst number, and labels causal explanations as guesses — all pressure against overstatement. The overstatement that remains is framing: a 99.5 percent failure rate for a category-leading engine is presented as an engine property when it was produced under an artificial 512 MB ceiling with no tuning attempt described, on a single run, in a test the vendor commissioned; the platform-count slip adds sloppiness to a piece whose whole argument is rigor.
Vendor-commissioned test, disclosed and partly self-penalising
The tester was engaged to benchmark CognoDB Cloud and chose the comparators under that brief, so the sponsor's product is structurally the reference point and the one engine that appears error-free across all concurrency levels. That is a significant incentive to distort. It is materially offset by disclosure of the commission, by publishing the sponsor's worst result in the entire test, by correcting a cap that would have disadvantaged the comparators, and by disabling a FalkorDB default that was suppressing a comparator's results. No compensation terms, editorial-control arrangements or vendor review of the draft are described.
Transparent but unreplicated single account
Confidence is limited by structure rather than candour: one publisher, one author, one run, self-reported numbers, no data release, and a commissioning relationship in the background. The clear methodology narrative, the retained adverse result and the explicit hedging of causal claims justify treating the reported measurements as plausible directional signals about behaviour under a 512 MB ceiling, but not as settled facts about any of these engines in production configurations.
build
A five-way graph database benchmark spent its first four hours measuring undersea cable1 distinct publisher
build
A CognoDB-side benchmark puts the vendor mid-pack, and ships the repo so you can check1 distinct publisher
invest
Prevalent AI takes $22m after nine years of self-funding, and points it at financial crime1 distinct publisher
invest
The balance sheet transfers, the relationship does not: $124tn and your donor file1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 21, 2026