Skip to content

Build1 publisher3 min readPublished

agentgateway v1.6.0 prices LLM calls from a catalog it ships itself

agentgateway v1.6.0 priced a Claude Sonnet 5 request at $0.000156 from its built-in catalog, with no hand-kept modelCatalog in the config. The test covered one Anthropic model, so retiring rate tables is safe only for models the catalog provably knows.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • In v1.5.0, per-key budgets only worked after someone hand-wrote a modelCatalog block listing every model with its per-million-token input and output rates.
  • A 503 caused by a dropped upstream TLS connection still used up a rate-limit slot, because localRateLimit counts requests the gateway admits.
  • The run was a standalone container, so the AgentgatewayModel CRD that the Helm chart now enables by default went untested.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Sizing a request bucket close to real traffic means upstream outages spend quota: one 503 consumed a third of the demo's per-minute allowance.
  • constraint Per-key limits multiply across replicas because localRateLimit buckets are per instance, so a fleet-wide ceiling requires moving to remoteRateLimit.
  • capability Operators can key quota on any CEL expression, such as a JWT claim or source IP, without adding a rule or restarting for each new key.

The logged cost checks out by hand. Eight input tokens at $2 per million plus 14 output tokens at $10 per million is $0.000156, the total agentgateway wrote to the access log [5][6][19]. Output tokens were $0.000140 of that, about 90 percent [20]. Those log lines come from one author's hands-on run against a live gateway and a real Claude Sonnet 5 backend, published with code and captured output in a demo repo [18].

In v1.5, the per-token rates had to sit in a modelCatalog block that someone kept current by hand [3]. In v1.6.0, generally available since October 2, they come from a catalog the project ships and maintains [1][4]. Getting them into the log still takes configuration. The author added a `frontendPolicies.accessLog.add` block mapping CEL expressions such as `llm.cost.total` and `llm.costRates.input` to log keys [7]. According to the post, a budgets policy reads those same fields, so the shipped catalog feeds the per-key USD budgets that arrived in v1.5.0 [8][2].

The demo config contains no budgets policy [9]. So the case that USD budgets now work without a rate table rests on that shared-field argument, with no overspend 429 observed. I think the argument holds. Before deleting a production table I would still want one budget tripped against catalog prices.

Coverage is the other open item. The post says the catalog covers more than Anthropic, but the run validated only claude-sonnet-5 [13][9]. It does not say how often the catalog changes between releases, or whether a hand-written modelCatalog entry overrides a shipped price. Retiring a table is therefore a per-model decision: a model whose log lines show populated `cost.rate.input` and `cost.rate.output` fields can lose its entry [6][7].

The rate limiter is good engineering. One localRateLimit rule keyed on `request.headers["x-api-key"]` gave key A three 200s and a 429 on its fourth request inside a minute [9][10]. Key B went through in that same minute, with no second rule and no restart [11][12]. Any CEL expression over the request can be the key, so a JWT claim or a source IP buckets the same way [12].

It counts admissions. In one run the upstream peer closed the connection without a TLS close_notify, the gateway returned 503, and `x-ratelimit-remaining` still dropped [14]. The provider produced nothing billable on that call, and the bucket charged for it anyway [14]. In a 3-request bucket, that one failure used a third of the key's minute [21]. The author wrote that "a bad backend day eats your quota exactly like a good one" [15]. The policy type is `requests`, so it caps calls, and spend control stays with budgets [9][2].

Buckets live in the proxy instance that created them. Quota shared across replicas is the job of a separate policy, remoteRateLimit [16]. Spread one key evenly across three replicas with this config and it gets up to nine requests a minute [22]. The run used a single container and no Kubernetes, so the AgentgatewayModel CRD, now on by default in the Helm chart, went untested [17].

What to watch

  • Whether agentgateway documents how often the built-in catalog is updated between releases and whether hand-written modelCatalog entries override shipped prices.
  • A published run that trips a per-key USD budget against catalog prices, and log checks for non-Anthropic models.
  • Kubernetes test runs of the AgentgatewayModel CRD now enabled by default in the Helm chart.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories