Skip to content

Build1 publisher3 min readPublished

Your agent traces are append-only, which is why they hide the bug

A write-up on assembling execution trees from span events argues the hard part is not collecting logs but handling out-of-order, concurrent, incomplete and retried spans without lying.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Your agent traces are append-only, which is why they hide the bug
Generated illustration

What happened

  • An agent trace is usually written as a sequence of events (span_started, span_ended) because append-only data is simple to produce.
  • Developers do not want to debug the raw event sequence; they want the causal structure, illustrated as a research_agent tree containing search_web, query_database, call_finance_api with attempt_1 timeout and attempt_2 ok, and summarize_results.
  • The event stream is optimized for writing; the execution tree is optimized for understanding.
  • Building a reliable tree requires more than sorting by timestamp: events may arrive out of order, siblings may run concurrently, spans may be incomplete, and retries may fail while the parent operation still succeeds.
  • Every event needs stable trace and span identity. Start events establish parentage; end events establish outcome and duration.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A dev.to write-up by Raju Dandigam takes apart a problem most agent observability stacks paper over: the trace format that is cheapest to emit is the one that hides causality. His argument is that agent traces get written as flat event sequences because append-only data is simple to produce [1], while the artefact an engineer actually needs is a causal tree showing an agent run branching into web search, a database query, a finance API call that timed out on the first attempt and succeeded on the second, and a summarisation step [2].

The framing is worth stealing: the event stream is optimised for writing, the execution tree is optimised for understanding [3]. The gap between them is not a rendering problem. According to the post, building a reliable tree requires more than sorting by timestamp, because events may arrive out of order, siblings may run concurrently, spans may be incomplete, and retries may fail while the parent operation still succeeds [4].

That last case is the one that quietly poisons dashboards. A tool call with a failed attempt_1 and a successful attempt_2 is a success at the parent and a failure at the leaf [2], and any pipeline that flattens status upward or downward will report the wrong thing.

The mechanics come down to identity. Every event needs stable trace and span identity, with start events establishing parentage and end events establishing outcome and duration [5]. Timestamps alone cannot do it: two adjacent events may be siblings, unrelated concurrent work, or operations from entirely different traces [6]. Nor can you assume the two halves of a span arrive together, since a buffered exporter can deliver the end before the start and a crashed process may never deliver an end at all [7].

Dandigam's assembler handles this by buffering an end event until its start arrives, keeping a separate map of pending ends [8]. Duplicate starts are dropped with a diagnostic rather than overwriting the existing span [9]. End timestamps are clamped with a max against the start time, so a skewed clock cannot produce a negative duration [10]. At finalisation, unmatched ends and spans still open are surfaced as diagnostics, and the assembler does not invent timestamps or mark incomplete work successful [11]. The diagnostic vocabulary is four codes: duplicate_start, duplicate_end, end_without_start, span_left_open [12] - two for redelivery, two for a missing half [13]. Six span kinds are declared: run, model, tool, retrieval, decision and fallback [14][15].

The post is honest about what this design does not solve. On an unbounded live stream, the pending-event map needs a size limit and an expiration policy, or malformed or hostile input can grow it without bound [16]. And a renderer should tolerate multiple roots and orphans without crashing, even though a valid trace normally has exactly one root [17].

Two things to watch in your own stack. First, whether your tracing UI can tell you that a span was left open, or whether it draws a tidy tree and omits the branch it could not resolve [11][12]. Second, whether the ingest path in front of it has a bounded buffer for orphaned ends [16]. A trace assembler that silently guesses is worse than flat logs, because it produces something that looks authoritative.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories