Skip to content

Build1 publisher3 min readPublished

No-key probes of four LLM gateways catch model lists changing within days

One developer's no-key test of four LLM gateways found model lists out of step with vendor docs, the author's own falling from 477 to 199 models in a week. Any model ID read from these lists is a dated observation and should be snapshotted before it goes into code.

The Engineer · Build desk

Illustration accompanying No-key probes of four LLM gateways catch model lists changing within days

What happened

  • The author called each list twice, twenty minutes apart, checking base URLs, field sets, counts, self-consistency and whether entries were safe to hard-code.
  • OpenRouter's public list returned 460 models on 29 September, against 445 in a snapshot taken ten days earlier.
  • In one case, a list returning id, context_length, max_output_tokens, capabilities and owned_by came back a week later with object, created, tier and modality and no capabilities field.
  • On the author's own site, the OpenAI base URL was right on the SDK page and wrong in the agent-facing file, both generated from one stale JSON registry.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost A base URL documented one path segment off fails as a 404 that looks like an auth error, so debugging time goes to credentials that were never the problem.
  • exposure Agents that configure themselves from a vendor's machine-readable file can receive a broken endpoint while people reading the SDK page get a working one.
  • constraint A published share such as tool-calling support holds only for the snapshot it came from, so any page quoting it needs that snapshot archived or a recompute at render time.
  • decision Hard-coding a model ID from a gateway list means keeping a dated copy of that list in the repository and re-checking it against the live endpoint before release.

An OpenAI-compatible client treats the base URL as the first half of a path. The Python SDK takes `base_url="https://example.com/v1"` and calls `/chat/completions` on top of it, and the `OPENAI_BASE_URL` environment variable works the same way, as the openai-python README documents [3]. An Anthropic client appends `/v1/messages` to whatever base it is handed [4]. If a vendor page gets its half wrong, every request built from it goes to a route that does not exist [5].

The check is one request, and the status code is the answer [7]:

curl -s -o /dev/null -w '%{http_code}' -X POST https://example.com/v1/chat/completions

A 401 means the route exists and wants a key. A 404 means the route does not exist, and no key will help [7]. The author recommends sending the same probe to `<base>/v1/messages`, because the two dialects are usually mounted at different prefixes and only one of them tends to be documented carefully [8]. I like this test. It needs no account and returns one number.

The author's own count drop, from 477 on 22 September to 199 on 29 September [9], is 278 entries, about 58% of the list in a week [1]. OpenRouter's rise was 15 models over ten days, roughly 3% [2]. The author's rule is that any figure read from an endpoint should carry the date it was read and the exact endpoint, so a reader can reproduce or refute it in one request [14]. "The fix is not to update the number more often. It is to stop putting a moving number in a sentence with no date on it," the author wrote [16].

Field drift costs the most. Once `capabilities` stopped arriving, every share built on it became uncomputable: tool calling, image input, combination routes [12]. The pages kept rendering and kept quoting the percentages, and nothing failed because nothing checked [12]. A percentage computed from a field that no longer exists does hold very steady. The author's first rule is to read `capabilities?.tool_calling`, which survives a missing key, and never `capabilities.tool_calling`, which crashes on the next release [13]. The other two are to avoid publishing a statistic derived from a field you do not control, and to snapshot the response into your own repository with its date [13]. In my view the snapshot rule matters more. Optional chaining turns a crash into an undefined value, and a share computed over undefined values is still wrong, with no error anywhere to say so.

The results the post reports are thinner than its headline, which says most gateways break the contract quietly [15]. Of the four gateways, OpenRouter is the only one named with a result [10]. The 477-to-199 drop and the base URL error were on the author's own site [9] [6], and the field change is described only as "the case I hit" [11]. The post does not report what the twenty-minute repeat calls returned [1]. The drifts it documents span seven and ten days [9] [10].

What to watch

  • Per-gateway results from the author for all five checks, especially the twenty-minute repeat calls and the gateways the post does not name.
  • Whether OpenRouter's next dated count keeps rising from the 460 read on 29 September.
  • Whether any gateway publishes a versioned schema or changelog for its model-list response.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories