Build1 publisher3 min readPublished
Wire-protocol fidelity is the first row on a gateway checklist published by someone who sells a gateway, and the tests that settle it cost nothing but curl. The rows that price your supply-chain risk need the vendor to answer in writing.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
A translation layer's losses live in fields, not in quality. The post names four casualties: tool-call block ordering, streamed tool arguments, prompt-cache markers and stop-reason fidelity, all of which live in fields with no equivalent on the other side [3]. A chat window reads none of those. An agent loop reads all four. That is why the first row asks which protocols an endpoint implements and which it translates, and treats those as two different products [2].
The max_tokens probe works because the Anthropic schema makes the field required. Send a `/v1/messages` call without it and an endpoint that validates the request itself has to reject it, which is the 400; an endpoint that returns 200 has supplied a value you did not send, on your behalf, before forwarding [4]. The 200 is the more informative outcome. A 400 tells you validation happens at the edge; it does not tell you whether something is rewritten between that edge and the model.
Streamed tool calls give you the same signal from the other direction. Incremental fragments carrying an `index` mean the endpoint is relaying deltas as they arrive; everything landing at once at the end means something buffered the upstream stream and re-emitted it [5]. Buffering is invisible in the final text and obvious to any consumer that acts on partial tool arguments.
Two more comparisons cost nothing. `GET /v1/models` should return the exact IDs your key can call rather than the vendor's full catalogue, and a completion response should echo a `model` field that matches what you asked for [7].
Then the availability number you generate instead of reading. A one-token completion against each model every five minutes is 12 calls an hour, 288 a day and 8,640 over a thirty-day month, appended to a JSONL file you own [18]. That buys a success rate and a latency distribution measured from your own network [12]. The author's framing of the alternative is the line worth keeping: a marketing-page availability figure is a claim, and a claim is not the same as a measurement [11].
One row is not about the vendor at all. Before a key is distributed, audit where it lands, because on a normal developer laptop that is an editor's settings store, a shell profile, two dotfiles, a CI secret and at least one Docker environment, five hiding places before the key has done any work [15].
The checklist is single-sourced and its author sells a gateway: he discloses working on daoxe, a multi-model gateway, and reports that two rows fail when he runs the list against it [16]. Treat that as a prompt for your own testing, not a settled finding. The tests cost curl calls, so run them against his endpoint, and against the one you are already paying for.
Ranked by verification strength, evidence, and original report placement.
The author discloses working on daoxe, a multi-model gateway, says he runs the checklist against it at the end of the piece, and states that two rows do not pass.
A dev.to post titled "Ten questions to answer before you route production traffic through someone else's LLM endpoint" presents a checklist of questions the author says a reader can verify themselves in an afternoon, arguing that the words on every LLM endpoint's landing page (fast, reliable, compatible, secure) are free.
The checklist's first question is which wire protocols an endpoint implements and which ones it translates into, and the post states these are different products.
The post says a translation layer is fine for text chat and lossy for agents, because tool-call block ordering, streamed tool arguments, prompt-cache markers and stop-reason fidelity all live in fields that have no equivalent on the other side.
Test given for the Anthropic protocol: omit the required max_tokens from a /v1/messages call. The post says a native implementation returns 400, while a translation layer often fills in a default and returns 200.
Test given for the OpenAI protocol: check whether streamed tool calls arrive as incremental fragments with an index, or all at once at the end.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Method you can rerun, results you never see
Most rows come with a test specific enough to fail: a malformed max_tokens request, a nonexistent model ID, a revoked key retried sixty seconds later. That is firmer ground than a vendor comparison table, because a reader with a key can falsify it. What keeps the score mid-range is that no test result appears anywhere in the piece, and the one claim about how aggregators actually source capacity rests on the author's say-so alone.
Nothing bought, deployed or measured
There is no release, benchmark, deployment or usage figure to weigh. The piece describes what to ask a provider before routing traffic, not what anyone has run or purchased, and the only product-level detail, daoxe's flat top-up rates across several payment rails, sits in a sentence that breaks off mid-clause.
Slightly ahead of what the text shows
The post argues against its own genre, calling a marketing availability figure a claim rather than a measurement and telling you to trust your own probe log instead, so the rhetoric mostly stays inside what the tests would prove. It still asks for a little credit it does not deliver: the promised score against the author's own gateway, two failing rows included, is not visible, and the passage that does surface about his employer is a payment-rails pitch.
Disclosed, and still shaping the list
The criteria elevated here, native protocol implementation, labelled capacity provenance, per-model key scoping, exportable usage, are the ones a well-built gateway passes and a thin reseller does not, and they are being proposed by someone who works on a gateway. The disclosure is in the second paragraph and unusually blunt about the conflict, and the author refuses to paraphrase any data policy including his employer's. Skewed rather than concealed, in other words.
One interested author, mostly checkable claims
Single publisher, single author, direct stake in the answer, and a text that truncates before its promised self-audit. What holds the score up rather than down is the nature of the assertions: procedural and self-verifying. If the max_tokens behaviour or the revocation lag is wrong, the first reader with an API key finds out.
product
Liability for a runaway agent lands on whoever configured its permissions1 publisher
product
Rillet's $100M reads as proof mid-market ERP is rip-and-replace, mostly at the cheap end1 publisher
invest
OpenAI rates GPT-6 Astra capable of hacking hardened systems without human guidance1 publisher
leadership
Disney swaps raises for discounted stock and a full health-plan re-enrollment1 publisher
Publishers with included, body-backed reporting in this cluster.
1 article · September 7, 2026