Build1 publisher3 min readPublished
The Cap Was One Line of Scheduler Policy, Not the GPU
A developer benchmarked his local agent runner before tuning it and found his instinct was wrong twice over: the event loop was idle, the hardware was half used, and the real work was serialized by a rule.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The author wanted to raise concurrency limits on his local AI agent runner; the UI now supports multiple terminal panes running in flight, and his instinct was that the runner process itself was becoming the bottleneck. He identifies himself as Andreas, a full-stack dev and CTO at a B2B SaaS, building his own agent tooling.
- Before touching a single config setting, he wrote a benchmark to test the assumption of what actually breaks when N sessions run at once: the model server, the hardware, or the scheduling policy. He states his gut was entirely wrong.
- The bench script fires N concurrent chat sessions against the runner. Each turn hits a local model (hermes3 via Ollama) with a fixed prompt, running the full pipeline: pre-turn intent classification, chat response, and post-reply memory extraction.
- Four metrics were tracked: ttfa (POST request to first internal event emitted, measuring the runner's event loop before touching Ollama), first-Ollama (POST to first request hitting the model queue), done (POST to user receiving the answer), and wall (drained) (total time until post-reply memory extraction finishes background work in Ollama).
- At N=4, wall-clock time scaled to 4.80x the single-session baseline; because 4.8x exceeds 4x, concurrent requests performed worse than a pure, perfectly ordered serial queue.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer building his own agent tooling wanted higher concurrency out of a local runner, assumed the runner process itself was the constraint, and wrote a benchmark before touching a single config value [1][2]. The measurement reassigned blame twice: chat throughput was pinned inside the local model server, and the coding jobs that do the actual work were serialized by one line of scheduler code [9][13].
The bench, described in Andreas Bodin's write-up on dev.to, fires N concurrent chat sessions at the runner, each turn hitting a local hermes3 model through Ollama with a fixed prompt and running the full pipeline: pre-turn intent classification, chat response, and post-reply memory extraction [3]. He instrumented four points rather than one: time to first internal activity, time to the first request reaching the model queue, time to the user's answer, and wall time until the background memory-extraction tail drains [4].
At N=4, wall time came in at 4.80 times the single-session baseline [5]. That is a 20 percent penalty against a perfectly ordered serial queue, which means four concurrent sessions were worse than doing the work one at a time [6][5]. Time to first activity stayed at 0.0s across every run, so the runner's event loop was routing POSTs in milliseconds without queuing [7]. CPU peaked at 63 percent and free RAM never fell below roughly 5.9 GB [8]. The latency was stacking in Ollama's internal request queue for the single loaded model, hermes3 8.0B Q4_0 [9]. The drained metric earned its place: background memory extraction keeps Ollama saturated well after the user has their answer, a cost that request-per-second benchmarks do not see [10].
Chat, though, is the cheap half. The coding jobs talk to the Claude API, not Ollama, and diagnosing those took no load test at all, only reading the scheduler [11][22]. The pump held a MAX_CONCURRENT value sourced from an environment variable with a default of 2, and immediately below it a check that skipped any job whose repository already had one running [12]. Since almost every ticket in the queue targets the same primary repository, the effective limit was one job, whatever MAX_CONCURRENT said [13]. The rule had a real justification, since two agents branching off the same moving HEAD produce messy merge conflicts, but serializing an entire repository by default is a policy choice, not a capability limit [14].
The replacement is a path-overlap predicate: two jobs in the same repo may run together only if both tickets declare target paths that share no files, ancestors, or subdirectories, neither references the other as a hard or soft dependency, and reopened tickets are judged by their previous run's actual git diff file list rather than their prose [15][16][17][18]. A ticket that names no paths is treated as touching everything and stays serial [16]. Ambiguity defaults to serial, on the stated asymmetry that an unnecessary serialization costs a few minutes of queue time while a bad parallelization costs an hour of untangling merges by hand [19].
What to watch: chat concurrency now binds on hardware or multiple inference instances, not TypeScript [20], and job concurrency binds on how honestly tickets declare their file scope [16][21]. No post-change throughput numbers are published yet.