Skip to content

Build1 publisher3 min readPublished

Your AI feature fails with a 200 OK, and your dashboard calls that healthy

A dev.to write-up argues per-request capture of prompt, model version, tokens, cost, latency and output quality is now a required line item, because crash-oriented telemetry has no field for any of it.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Traditional monitoring rests on an unwritten assumption that the same input gives the same output: something breaks, you replay the request, you watch it break again, you fix it.
  • Send the same request to a model twice and you get two different answers, and neither one threw an error.
  • AI observability is defined as recording what happened inside an AI system on every request: the prompt, the model version, tokens, cost, latency, tool calls, and a judgement of whether the output was any good.
  • Monitoring tells you the service is up; observability tells you why it answered that way.
  • Existing monitoring setups watch for crashes using status codes, error rates, p99 latency and memory, all designed around the idea that a broken thing looks broken.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

The failure that matters in an AI feature does not look like a failure. In a dev.to explainer, an engineer who built a retrieval chatbot at Keploy describes the shape precisely: HTTP 200 in 900ms, grammatically perfect prose that happens to be wrong, or that ignored the document you just retrieved for it, or that called the refund tool when the user only asked a question [6][10].

The reason your stack misses this is structural, not a tuning problem. Conventional monitoring is built around status codes, error rates, p99 latency and memory, on the assumption that a broken thing looks broken [5]. By every measure the dashboard has, the service is healthy [7]. Underneath sits an older assumption that nobody writes down: the same input gives the same output, so you replay the request and watch it break again [1]. Send the same request to a model twice and you get two different answers, neither of which threw an error [2].

The author's definition of AI observability is a list of per-request fields: the prompt, the model version, tokens, cost, latency, tool calls, and a judgement of whether the output was any good [3]. Compare that list against the crash-oriented one and the overlap is zero [19]. There is nowhere in standard telemetry to put "this response cost 14 cents", "the model version changed under us last Tuesday", or "the retrieved context was garbage", because those are not infrastructure facts [8].

The mechanics are unglamorous. One question to a retrieval system is a chain, not an operation, and a trace is that chain written down as spans, each with its own start time, end time, inputs and outputs [11]. The question opens a root span; embedding it into a vector is a span with its own model and cost; the vector search is a span whose important content is which five chunks came back; the model call carries model name and version, temperature, prompt and completion tokens, cost in dollars, total latency and time to first token; each tool call is a child span [12].

That structure buys one specific thing: the difference between two fixes. If the search returned five irrelevant chunks, the problem is chunking or embeddings and the model did nothing wrong; if it returned exactly the right documentation and the model still answered from thin air, the problem is the prompt [13]. Without the trace you have "the bot said something dumb", which the author calls the most useless bug report there is [13].

Notably, the piece drops the "three pillars" framing on purpose, arguing that logs, metrics and traces was designed for deterministic systems and has no slot for whether the answer was good [14]. What replaces it is a field list: the final assembled prompt and response as they went over the wire, not the template [15]; model, version and parameters, because providers ship silent updates and you cannot explain last month's regression without knowing which version answered [16]; tokens in and out plus dollar cost attributed to a user or feature [17]; and latency split into total time and time to first token, which feel like different products to a user [18].

Watch the version field first, since silent provider updates are the failure you cannot reconstruct after the fact [16]. Also note the disclosure: the author's team uses Bifrost, an open-source AI gateway from Maxim, and the argument for it is that per-request recording claims can be checked in the code rather than on a marketing page [9].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories