Build1 distinct publisher2 min readPublished
CTGT's 76-prompt audit puts an anonymous OpenRouter endpoint at 6.5 on the usual censorship probes and 89.5 on Chinese domestic legitimacy. The prompt list is the benchmark.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Average the seven topic pairs that carry Ox Alpha's measured gap and the separation between a sensitive prompt and its matched control comes to 79.2 points [2]. The floor of that set is the 2018 removal of presidential term limits at 61.25, with the domestic incidents just above it: the Xuzhou chained-woman case at 80.5, the Sichuan school collapses at 73.5, the Wukan protests at 73 [6]. Nothing in that spread reads as a model weighing sensitivity. It reads as a list.
The distribution agrees. Nothing at all landed in the 25 to 50 band [5]. Ten of the 76 responses fall outside both of the clusters CTGT described [1], and those ten are the part of the audit an outside reader cannot reconstruct from the thread.
The comparison that matters is the model against itself. On seven Xinjiang and Taiwan prompts, which is where standard audits start because those subjects have historically produced refusals or state-aligned answers from Chinese models [10], Ox Alpha averaged 6.5, a point under OpenAI's GPT-OSS-120B [7][3]. On four prompts about Xi and Communist Party legitimacy it averaged 89.5 [8]. Same endpoint, frozen decoding [4], 83 points apart, a factor of nearly 14 [4]. DeepSeek V4 Flash moves 25.2 points across the same two buckets [5], which is what a broad censorship profile looks like when you plot it: high everywhere and flat. Ox Alpha sits 2.2 points above DeepSeek on Xi and legitimacy and 5.4 below it on domestic protest and accountability [8].
Fourteen prompts out of 76 do all of that work [6]. The other 62 are what a short checklist samples, and they answer.
On identity, both investigations are measuring a vocabulary rather than a politics. CTGT matched live token counts exactly against GLM-5.2 across all 11 of its aggregate probes, with GLM-4.6 matching nine and Kimi, MiniMax, DeepSeek, Qwen, OpenAI and Microsoft-family candidates producing larger errors [11]. The independent black-box run published on GitHub reported a 44-for-44 tokenizer match with the GLM-5 generation over more than 600 requests and roughly 13.5 million prompt tokens, alongside GLM-style reasoning controls and Chinese-language gateway errors [12]. Strong evidence about the tokenizer is weak evidence about the filter, because a gateway moderation layer can shape answers independently of the weights being served [14].
Which leaves the operational reading. A refusal rate for this endpoint is a statement about somebody's prompt list, and Gorlla's seven pairs decide the number [1]. Omit them and the same endpoint scores 6.5 [7].
Ranked by verification strength, evidence, and original report placement.
Cyril Gorlla, founder and CEO of AI control startup CTGT, found that the anonymous Ox Alpha model sharply restricts answers about seven threats to Chinese domestic political legitimacy while responding freely to several subjects that usually expose censorship in Chinese models.
Ox Alpha appeared on OpenRouter on August 20th without a named developer.
The two investigations do not establish who operates Ox Alpha or which exact checkpoint is served: CTGT's compact fingerprint matched GLM-5.2 while the independent investigation ranked the original GLM-5 as its best checkpoint fit, and both treat the GLM-family attribution as stronger than the version-level attribution.
A gateway moderation layer can shape answers independently of the underlying weights.
OpenRouter's listing describes Ox Alpha as a model developed and operated by an anonymous third party, presents it as a free reasoning model for coding, and OpenRouter says it only routes requests and is not the developer, owner or operator.
Gorlla published the results in a six-post thread on X on Tuesday.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Disclosed method, single-auditor scores, corroborated fingerprint
The audit design is stated (76 China-sensitive prompts, matched controls, frozen decoding, four AI judges) and results are reported at per-topic granularity, which is stronger than a bare claim. But the censorship scores come from one vendor-run test relayed by one publisher, with no published prompt list, control wording or judge identities, and only 14 of 76 prompts broken out. The GLM-family attribution is materially better evidenced because two independent probes converge on it, while the article itself limits both operator and checkpoint conclusions.
Publicly routable and actively probed, no usage data
There is concrete deployment evidence: the endpoint has been live on OpenRouter since August 20th as a free reasoning model with a 1.05M-token context window, and it has drawn at least two independent probing efforts, one of which issued 600+ requests. But no usage, traffic, customer or revenue disclosure appears anywhere in the supplied material, and the operator is anonymous, so real-world dependence on this endpoint cannot be sized.
Hedged framing, but precise single-source numbers
The article is unusually restrained for this topic: it declines to name an operator, flags the GLM-5.2 versus GLM-5 divergence, offers gateway moderation as an alternative explanation, and says Chinese AI rules do not prove Z.ai involvement. The mild overstatement comes from decimal-level precision (94, 93.75, 89.5, 6.5) carried from one unreplicated vendor audit whose prompts and judges are not published, and from bucket means built on seven, four and three prompts being used to characterize an endpoint's overall policy.
Vendor-run audit doubling as product proof, disclosed
The measurement comes from CTGT, a venture-backed startup (Gradient, General Catalyst, Y Combinator, Liquid 2) whose business is measuring and controlling model behavior, published by its founder and CEO on X; publishing striking censorship findings advances that commercial position. The article discloses this affiliation and funding, which mitigates but does not remove the incentive. Counterweights are limited: OpenRouter's role is a disclaimer, the endpoint operator is anonymous with unknown motives, and the corroborating GitHub investigation's authorship and incentives are not described.
Family attribution firm, scores single-sourced
Confidence splits by claim type. That an anonymous GLM-family-fingerprinted endpoint is live on OpenRouter is well supported and independently corroborated. That its censorship is narrowly concentrated on domestic legitimacy topics rests on one vendor audit, relayed by one publisher, without a published prompt set or replication, and the version-level attribution is openly divergent between the two investigations. Adoption scale is unknown and no operator has responded.
science
GLM-5.3 says the quiet part: the base model did not change, the post-training did1 distinct publisher
build
GLM-5.3 is a paper, not an endpoint: Z.ai publishes research before weights1 distinct publisher
build
WRITER's new flagship is a post-train of Z.ai's GLM-5.2, and that is the story1 distinct publisher
invest
DeepSeek V4 doubled its OpenRouter token share, and the bill it displaced was ~130x bigger1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 25, 2026