Skip to content

Build1 publisher3 min readPublished

Stop timing your GraphQL tests and start counting loader calls

A dev.to post argues latency assertions cannot catch an N+1 because dev machines are too fast. The proposed gate is a batch-size log and a call count that must stay flat.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • You will catch a GraphQL N+1 much earlier by counting how many times your batch loader is called against a temporary mock server than by measuring response time.
  • Latency tests hide the problem behind small local datasets, in-memory caches, and a developer machine that is simply too fast to notice.
  • The classic mistake is to add a deeply nested field, run one query that returns three posts, see a snappy response, and ship.
  • That query may call the author loader three times instead of once, but each call is so cheap that nothing complains until the same shape runs in production with fifty posts, remote storage, and a cold cache.
  • The rule is not about rendering the graph or optimizing resolvers; it is about observing a number that should stay flat.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A post on dev.to argues that the earliest reliable signal for a GraphQL N+1 is the number of times your batch loader gets called against a temporary mock server, not the response time of the query [1]. That matters because the gate most teams already have, a latency assertion in CI, is defeated by the conditions the test runs under: small local datasets, in-memory caches, and a developer machine that is simply too fast to notice [2]. The failure mode is dull and familiar. You add a deeply nested field, run one query that returns three posts, see a snappy response, and ship [3]. That query may call the author loader three times instead of once, but each call is cheap enough that nothing complains until the same shape runs in production against fifty posts, remote storage, and a cold cache [4]. The alternative is to assert on a number that should stay flat [5]. If a post has an author and an author has posts, a query returning ten posts and then asking for each author should produce exactly one author batch of size ten, not ten batches of size one [6]. That distinction is a tenfold difference in loader invocations for identical data, which is why it survives being measured in milliseconds on a laptop and does not survive being counted [23]. If the mock server records the size of every batch and fails when the pattern is violated, you get a regression gate that does not depend on the data model, the network, or anyone's intuition about what is fast enough [7]. The fixture is small on purpose. Authors are treated as a remote resource; `get_author_batch` receives a list of keys, appends the length of that list to a module-level log, and returns fake records [8]. The schema needs a `Post` type, an `Author` type, and a nested `posts` field on author so a query can loop back through the same resolver [9]. ```python batch_sizes = [] def get_author_batch(keys): batch_sizes.append(len(keys)) return [{"id": key, "name": f"author-{key}"} for key in keys] ``` That is the whole instrument [10]. The assertion is: after executing a query, take the largest recorded batch for that loader and compare it against the number of unique keys in the response. Ten distinct authors means the largest batch should be ten and the total call count should be one; ten calls of size one is the bug, found before anything is slow [11]. The article also pitches using a free model to generate the query corpus, and discloses that it was prepared as part of MonkeyCode's product outreach [12]. The narrow version of that suggestion is defensible: hand the model the schema and ask for documents that traverse the same relationship through aliases, fragments that re-include the author, pagination arguments that change list length, and cycles two or three levels deep [13]. The model writes GraphQL strings; it does not judge correctness, because your existing runner executes them [14]. The vendor claim is that free model access suffices for this because the task is narrow and the output is easy to review, and that a free disposable HTTP host can run the counting mock [15]. Treat the second half of that as marketing; the counting mock runs anywhere. The interesting cases are the ones a human test author skips [16]. A fragment can expand the same author through two aliases, and a loader that keys on alias rather than identity will be called once per alias even though the underlying keys are identical [17]. A `posts { author { posts { author } } }` query forces the batching strategy to re-enter the same loader before the first batch has resolved [18]. In CI, this is a loop: start the server once for the test file, execute the corpus, and read a per-request batch log off an endpoint [19][20].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories