Skip to content

Build1 publisher3 min readPublished Updated

Per-tenant Claude clients belong in the dependency graph, not in middleware

A FastAPI writeup on isolating Anthropic keys per tenant is really a bug report about ambient request context, and about what breaks the moment a background task reads it.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The author writes that CitizenApp reached 15 tenants while running a single global Claude API key, which the author called a ticking time bomb.
  • The author states that one customer's agentic loop burning through their quota would throttle everyone else, and that there was no way to enforce per-tenant rate limits without middleware spaghetti that would make debugging a nightmare.
  • The stated fix is to use FastAPI's dependency injection system to make tenant-specific Claude clients and rate-limit buckets first-class citizens, with "no globals, no thread locks" and no detective work about who is using the API key.
  • The team previously had a get_current_tenant() middleware that set request.state.tenant_id.
  • With that middleware in place, handlers had to manually fetch the API key and pass it around.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A post on dev.to describes CitizenApp arriving at 15 tenants with one global Claude API key, and the author's conclusion that a single customer's agentic loop could burn the shared quota and throttle everyone else [1][2]. The interesting part is not the quota arithmetic but where the author says the previous design actually failed: background tasks that queried Claude after the request had moved on [7].

The earlier arrangement will look familiar. A `get_current_tenant()` middleware set `request.state.tenant_id`, and handlers then fetched the API key themselves and passed it down the call chain [5][6]. According to the author, once background tasks started calling Claude, context vars leaked, rate limits went unenforced, and working out which tenant a given call belonged to took hours [7]. That is the predictable end state of ambient context: middleware runs once per request, so the tenant's key has to be parked somewhere a later reader can find it, and shared rate-limit buckets get touched on the hope that concurrent requests do not collide [8][9].

`Depends` changes the shape of the problem rather than the amount of code. Dependencies resolve per request, compose into other dependencies, and can be cached per dependency with `use_cache=True` [10]. The practical difference is that a handler receives a client object it can hand to a background task explicitly, instead of a task reaching back for request state that may no longer be the right one [4].

The implementation is small. Tenant rows carry the Anthropic key, which the author notes should be encrypted in production, a `max_requests_per_minute` defaulting to 60, and a `preferred_model` defaulting to claude-3-5-sonnet-20241022 [11]. `get_tenant_id` decodes a bearer JWT and raises 401 on a missing or invalid token [12]; `get_tenant` loads the row once per request and raises 404 [13]; `get_claude_client` and `get_rate_limit_bucket` are kept as separate providers so either can be injected without the other [14]. Note that this puts a database lookup in front of every model call, and the per-dependency cache only deduplicates within a single request [13][10].

Two things in the listing deserve flagging. The post opens with "no globals", but the clients and buckets live in module-level dicts keyed by tenant id [3][15][3]. And the buckets are in-memory sliding windows, which the source itself says should be Redis for distributed deployments [16]: run four workers and the enforced ceiling per tenant becomes 4 x 60 = 240 requests a minute, not 60 [1]. The client cache has a subtler edge. `get_claude_client` only constructs a client when the tenant id is absent from the dict, so a key rotated in the database is not observed by a warm process until that process restarts [17][2].

What to watch: whether the bucket store moves to Redis before the worker count changes, because the limit multiplies silently [1]; whether the per-tenant client dict acquires an invalidation path for key rotation and tenant offboarding [2]; and whether background tasks are actually handed the injected client as an argument, since that is the exact case that broke the middleware version [7].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories