Skip to content

Build1 publisher2 min readPublished

LiteLLM's Lens moves agent-failure analysis into the self-hosted gateway that routes model calls

LiteLLM launched Lens on September 30, a tool that uses AI agents to find recurring failures across agent traces sent through its model gateway. Customers host the analyzer and its databases, and the 200,000-trace volume CTO Ishaan Jaffer cites is a future target Lens has not been measured against.

The Engineer · Build desk

Illustration accompanying LiteLLM's Lens moves agent-failure analysis into the self-hosted gateway that routes model calls

What happened

  • Teams send agent traces to the LiteLLM proxy over OpenTelemetry, define what a good run looks like, and have Lens check selected runs against criteria such as recovery from a tool failure.
  • Jaffer's launch thread on X pitched agent-first APIs, SQL queries over traces, and letting Codex or Claude Code analyze them without a hosted platform's rate limits.
  • Lens is available only through an early-access waitlist.
  • LangChain added a chronological agent-session view to LangSmith on September 24, six days before LiteLLM's Lens post.
  • Jaffer previously co-founded Berri, a Carnegie Mellon alumni startup in the 2023 VentureBridge cohort whose product watched LLM applications for hallucinations and refusals.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Platform teams that adopt Lens take on three more components to run beside the proxy, and every investigation adds model calls on a model they choose and host the analyzer for.
  • capability Security teams can review what agents did without shipping trace data to a separate hosted observability service.
  • constraint Lens surfaces only the failures a team has instrumented and written checks for; it does not make agents correct themselves.
  • decision Langfuse can also be self-hosted, so LiteLLM users are really deciding whether one proxy holding routing, spend and agent records is worth giving up a dedicated observability tool.

A trace records one run: the model calls, the tool use and the intermediate steps between them [11]. When an agent fails, the cause can sit several steps before the bad output, so a single request-response record often does not show it [11]. Lens works across runs. It groups similar problems and links each group back to the traces that produced it [6]. Investigations run on demand or on a schedule, and the individual traces stay available in the proxy's logs interface for manual inspection [10].

I think the storage design is the right one for a self-hosted product. The step-by-step trace record goes to ClickHouse, and the investigation results that point into it go to PostgreSQL [7]. A separate analyzer worker connects to the proxy and runs on the customer's infrastructure [7][8]. The analyzer is also a gateway client. It calls a model the customer selects, routed through LiteLLM [9]. Every investigation therefore adds model traffic to the proxy [7][9].

The rate-limit pitch follows from where the data lives. LiteLLM's stated benefit is direct access to customer-controlled trace data without the rate limits of hosted tracing platforms [24]. A coding agent running SQL against that ClickHouse instance is limited by the hardware the customer gave it [7][24].

Jaffer said Lens is being built for a future in which agent swarms generate "200K+ traces" [4]. The launch image illustrates it with hundreds of agents making thousands of model and tool calls, each run a trace carrying steps, tokens and cost [15]. The figure is a design target [4]. For it to tell an operator anything, LiteLLM would need to publish an ingest rate into ClickHouse, how many runs the analyzer can investigate per hour, and what one investigation costs in model calls. The documentation does not establish general availability, Lens pricing or measured performance at the announced scale [14].

Much of the feature list already exists elsewhere. LangSmith offers trace monitoring with automatic clustering and failure analysis [18]. Arize Phoenix and Braintrust combine tracing, monitoring and evaluation [20]. LiteLLM's claim to difference is position. Jaffer described the gateway as a chokepoint for enterprise AI traffic [3]. Few vendors pick that word for their own product, and Runtimewire treats it as company positioning [3]. Runtimewire also notes a commercial incentive in the expansion: the open-source gateway is free to self-host, while the enterprise offering adds governance, security and support [21]. LiteLLM is a Y Combinator Winter 2023 company and raised a $1.6 million seed round from Y Combinator, Gravity Fund and Pioneer Fund [23].

What to watch

  • Throughput and per-investigation cost figures from LiteLLM or an early-access customer running Lens near the "200K+ traces" volume Jaffer described.
  • Lens pricing, and whether it ships in the free self-hosted gateway or only in the enterprise tier.
  • Whether teams already on LangSmith or Langfuse move trace analysis into the LiteLLM proxy once the waitlist opens.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories