Skip to content

Build1 publisher3 min readPublished

Fail-open plugin left a self-hosted Langfuse tracing 7% of one developer's agent calls

One developer's self-hosted Langfuse traced 91 of 1,387 agent model calls in 14 days because only one of ten profiles had its keys. The plugin fails open by design, so nine untraced profiles, one a coding agent with more than five times the traced calls, raised no error.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Fail-open plugin left a self-hosted Langfuse tracing 7% of one developer's agent calls
Generated illustration

What happened

  • A Discord approval to post a drafted dev.to reply ran 25 minutes, posted nothing and returned a confident summary of the article that was made up.
  • Jobs launched in safe mode skipped plugins entirely, so the reply drafter went untraced until it was switched to the Langfuse SDK.
  • ClickHouse, the Langfuse v3 trace store, used 6.03 GiB of disk to hold 2.2 MiB of traces, with the rest taken by its own diagnostic logs.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure The coding agent made more than 455 calls with no timing record, so a slow turn there would be diagnosed by feel, the error the one trace prevented.
  • constraint Per-profile key scoping plus a silent fail-open means every new profile starts untraced, and only an outside count, such as the session database, will show it.
  • cost Langfuse v3 cannot run without ClickHouse, so a small install on a shared 5400 rpm disk pays for continuous log writes to store very little trace data.

According to the author's dev.to post, the three model calls in that trace add up to 1,517.0 seconds. The three tool calls add up to 4.6 [2][3]. Tools took 0.3% of a 1,526-second turn [1]. Without the trace, the author wrote, the web tools would have taken the blame [4]. The post puts it plainly: "The trace says the network was the fastest thing in the room." [5]

The timings ruled out the tools, but they did not contain the cause. The Discord conversation was one long-lived session, started days before the reply skill existed, and the agent reads its skill list once, at session start [6]. Asked to post with no posting tool, the model searched for the article, extracted the page and wrote a plausible summary [7]. Finding that took knowledge of how the agent loads skills. A fresh session fixed it [8].

The coverage gap comes from a sensible design. The tracing plugin is enabled per profile. It reads its Langfuse API keys through that profile's own secret scope, so profile B's traces never ship under profile A's keys [14]. The author added keys to the default profile only [15]. Without keys the plugin's hooks do nothing, and the specialist profiles ran untraced with no error and no log warning [15][16]. The author calls failing open the right choice for a plugin [16]. I agree; a tracing outage should not stop an agent. "A monitoring tool that's switched off looks identical to one that has nothing to report," the author wrote [17].

The gap showed up in the counts. Langfuse held 618 events across 79 traces in four weeks, about three traces a day [10]. The agent's own session database logged 188 sessions and 1,387 model calls across ten profiles in the last 14 days [11]. The traced default profile had 91 of them [12]. That leaves 1,296 calls, about 93%, untraced [3]. The coding agent alone made more than five times the traced profile's calls, so more than 455 [13][4]. Safe mode opened a second hole. One-shot jobs launched that way skip plugins entirely, and the reply drafter used it for speed [18]. The drafter now sends its own traces through the Langfuse SDK [18].

The post does not report a misdiagnosis on any untraced profile. What it shows is narrower. The one traced profile got an elimination step, and the profiles carrying most of the calls had nothing to check a guess against [3]. Blaming the model by reflex would mislead too. In the traced data the slowest single tool call was a search_files at 120.6 seconds, while terminal ran 135 times at a 0.6-second median [23].

The same data puts the local model, on a 6 GB card, at a 125-second median, about 30 times slower than the fastest cloud option [9]. That figure is one card running one model on one person's jobs. It transfers to another setup only if the hardware and the workload match. It also comes only from the default chat profile [9][12].

Self-hosting had its own cost. From v3, Langfuse stores traces in ClickHouse, with Postgres for users and projects, Redis for the queue and MinIO for blobs, so ClickHouse cannot be dropped [19]. Here it used 6.03 GiB of disk to hold 2.2 MiB of trace data [20]. The rest was ClickHouse's own diagnostic logging, roughly 2,700 times more bytes about ClickHouse than about the agents [20]. Those tables are written continuously. The node keeps its guests on a single 5400 rpm HDD that already has a 15-minute IO storm after every reboot [21]. "So the trace store I'd barely been using was grinding that disk all day, recording its own profiler samples," the author wrote [24].

The coverage fix keeps the scoping. For each of the ten profiles, the author copied the Langfuse keys into that profile's own env file [22].

What to watch

  • Whether the agent's tracing plugin adds a startup warning when a profile loads it without Langfuse keys.
  • Per-model latency once all ten profiles report, especially how the coding agent splits work between the local and cloud models.
  • How the author trims ClickHouse's diagnostic log tables, and what disk use looks like afterwards.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories