Skip to content

Build1 publisher3 min readPublished Updated

Flat logs cannot explain an agent run, and adding more of them makes it worse

A dev.to walkthrough argues the fix is not abandoning console.log but giving events identity and parentage: traceId, spanId, parentSpanId, one start, one completion.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Flat logs cannot explain an agent run, and adding more of them makes it worse
Generated illustration

What happened

  • A dev.to article by Raju Dandigam, "Why console.log Isn't Enough When Building AI Agents", argues that the limitation in debugging agents is not console.log() itself but flat, uncorrelated events.
  • The article states agent debugging needs identity, parent-child relationships, lifecycle and safe metadata, and that without those, more log lines often create more noise rather than more understanding.
  • The article's sample flat output reads: 10:00:01 search started; 10:00:01 search started; 10:00:02 model started; 10:00:02 search completed; 10:00:03 search timed out; 10:00:03 cache fallback used; 10:00:04 model completed.
  • The article lists questions the sample output leaves open: which search completed and which timed out; whether the searches were siblings or one was a retry; which model call depended on which search result; whether the model began before retrieval finished; whether the final answer used live or cached data; and which user request produced the events.
  • The article states that timestamps describe when events were written and do not describe causality.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A post on dev.to by Raju Dandigam makes a narrow, useful argument: the thing that stops explaining an agent run is not `console.log` but flat, uncorrelated events [1]. It matters because the usual response to a confusing failure is to add log lines, and the author's claim is that without identity, parent-child relationships, lifecycle and safe metadata, more lines produce more noise rather than more understanding [2].

The demonstration is a seven-line terminal dump covering three seconds of wall clock: two identical `search started` lines at 10:00:01, a model start, a search completion, a search timeout, a cache fallback, and a model completion [3][14]. Every event is present and the run is still unreadable. According to the post, you cannot tell which search completed and which timed out, whether the two searches were siblings or one was a retry, which model call depended on which search, whether the model started before retrieval finished, whether the final answer used live or cached data, or which user request produced any of it [4]. Timestamps record when a line was written, not what caused it [5].

Short sequential request paths are fine with flat logs, the author concedes [6]. Agents are harder because control flow is decided at runtime: a model picks the tool, several retrieval strategies run concurrently, a failed tool is retried with different arguments, a fallback returns stale but valid data, one agent hands work to another, and a stream starts before the complete result or token usage is known [7]. The proposed representation is a tree, where the two model calls sit under `classify_question` and `generate_answer` and therefore have visibly different roles, the parallel searches are children of `retrieve_context`, and the cache fallback belongs to `check_account` [8].

The strongest part of the argument is about failures that never throw. A workflow can report success end to end while taking the wrong path [9]. The example is a quote agent that says an item is available: every top-level operation reports success, the HTTP status is 200, the response is syntactically valid, and the inventory number came from a cache 24 hours old because the live service timed out [10]. In tree form that is legible at a glance, because the timeout sits under `check_inventory`, the cached child carries `age_hours=24`, and `compose_quote` records `inventory_source=cache` [11].

Then the honest concession: flat logs can carry all of this, but only if every line holds enough context to rebuild the relationships, at which point you have already built a tracing model [12]. The suggested contract is a ten-field event type, three fields of which are pure identity: `traceId`, `spanId`, `parentSpanId`, plus event, name, kind, timestamp, and optional status, duration and metadata [13][19]. The writer is still `console.log(JSON.stringify(event))` [14]. Consumers group by `traceId` and rebuild parentage from `parentSpanId` [15], and filtering stops being a prose grep and becomes a query on `kind=tool` and `status=error` [16].

What to watch is discipline, not tooling. The model rests on one start and one completion per meaningful span, with status and duration on the completion [17][18]. Streaming is the case where that will strain, since the post itself notes the stream opens before the result or usage is known [7], and it is where partially written spans and orphaned children will show up first.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories