Invest1 distinct publisher3 min readPublished
Accuracy scores stay flat while response behaviour falls apart. The fix Tencent proposes costs under a point of accuracy, but it lives at training time, where the operator routing for cost has no access.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
The routing rule that produces this is usually set once, in a config file, by somebody weighing per-token price and latency across two modes of the same checkpoint, and then left alone, because the suite that would reopen the question scores answers against ground truth, and the argument in the Tencent paper is that the damage sits somewhere ground truth never reaches [1][4].
So take the two numbers the paper actually publishes and add them: chain-of-thought leakage at roughly 12.75% and response repetition at roughly 8.49% [6] give a combined average trigger rate of 21.24% [12], which is about one response in every five [19], and on a 2,415-prompt benchmark [2] the leakage figure alone works out at something like 308 prompts [16]. Now set 21.24 against the flagship gap of 48.64% [3]: the two patterns with disclosed averages account for 43.7% of it [13], which leaves logical contradiction and performative reasoning [5] carrying the remainder, assuming the rates are additive, which the summary I have does not establish [18]. Or rather, the more useful version of that observation: the two failure modes nobody has put a rate on are the two that a downstream parser will happily ingest.
The remedy arithmetic is the cleanest thing in the paper. PatternRL cut non-thinking trigger rates by 13.08 percentage points on Qwen3-VL-4B and 14.35 on the 8B, with accuracy moving less than a point either way [9][10], so the exchange rate is at least thirteen points of response behaviour for one point of accuracy [15], and 13.08 removes about 61.6% of the 21.24 floor [14]. The catch is where the lever sits: PatternRL writes pattern-specific penalties into reinforcement learning during fine-tuning [8], which an operator calling a hosted endpoint cannot do. What they can build is the other half, PatternRM, a response-level reward model that scores the four patterns [7], and that is an eval line item, not a training run.
That headline number deserves some caveats before it becomes a verdict. The 48.64 is a ceiling, reported as "as high as" in flagship models [3], so a team on a smaller model may be living nearer the 12.75 and 8.49 averages [6]. Benchmark prompts are not production traffic, and PatternEval is explicitly a diagnostic instrument [2], so the trigger rates need not transfer to a narrow workload. And the vendors may simply fix this upstream, which the 8B improving 9.7% more than the 4B [17] mildly supports: alignment moved the number, and it moved it more with scale.
This is probably wrong, but the honest read is that non-thinking mode is being sold as a cost lever and consumed as a quality decision, with the quality half unpriced because the dashboard was built to grade answers. What would kill the thesis: a team instrumenting the four patterns on its own traffic and finding trigger rates in the low single digits, at which point the 21.24 is a benchmark artefact and the cheap route was simply cheap.
Ranked by verification strength, evidence, and original report placement.
Tencent researchers published a paper on arXiv titled "Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs".
PatternEval is a diagnostic benchmark of 2,415 multimodal prompts spread across multiple task categories, which evaluates the quality and coherence of the response itself rather than only whether the answer was correct.
The gap between thinking and non-thinking inference failure rates reaches as high as 48.64% in flagship models.
Hybrid-thinking models do not necessarily get things wrong more often in non-thinking mode; accuracy numbers can look comparable on a benchmark dashboard while the actual experience diverges sharply.
The researchers identified four dominant failure patterns in non-thinking outputs: chain-of-thought leakage, response repetition, logical contradiction, and performative reasoning.
Average trigger rates were roughly 12.75% for chain-of-thought leakage and 8.49% for response repetition.
Distinct publishers with included, body-backed reporting in this cluster.
cryptobriefing.com
1 article · August 30, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
1.5% of Hugging Face repos take 99.2% of downloads, and the ceiling is Chinese1 distinct publisher
build
Ten agent eval protocols say when a run stops. Fewer say whether the result is settled.1 distinct publisher
invest
XPeng's $900m robot spin-out comes with a seven-year clock and a $1bn put2 distinct publishers
build
China's accelerator swap makes Cambricon supply, not export policy, your ship-date risk1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Sharp decimals, one hand
The 48.64-point gap, the 12.75% and 8.49% trigger rates and both Qwen3-VL reductions all arrive via a single retelling on cryptobriefing.com of a Tencent preprint. The paper's tables are never quoted, the flagship models behind the headline number are never named, two of the four failure patterns carry no rate at all, and no one outside Tencent has run the benchmark. Three-significant-figure precision is doing work that verification has not done.
Trending list, nothing shipped
What actually exists: a preprint, a benchmark no third party has run, and a training recipe demonstrated on two small Qwen3-VL checkpoints inside the authors' own experiments. The external signal Crypto Briefing offers is a Hugging Face top-paper placement, which measures reading, not deployment. No vendor has said it will penalise chain-of-thought leakage in a shipping model, and this coverage does not establish that the benchmark or the reward model is even available to download.
Maximum in the headline, average in the body
Crypto Briefing's headline promises failures "up to 48%" and converts a percentage-point spread into a percentage increase — two different claims, and the larger one runs first. Read down and the piece behaves better: it says plainly that non-thinking mode is not necessarily less accurate, which is the genuinely interesting result, and it reports the sub-one-point accuracy cost without embellishment. The inflation sits at the top of the story, not throughout it, and the word "promising" is at least labelled as early.
Same lab names the disease and sells the cure
Tencent defines the four failure patterns, builds the benchmark that measures them, then supplies both remedies and scores them on that same benchmark. The demonstration runs on Alibaba's Qwen3-VL weights rather than Tencent's own flagship, and the models producing the unflattering 48.64 figure go unnamed — an arrangement that spares every competitor and tells the reader nothing about where the risk actually lives. None of this makes the numbers wrong; it does mean the ruler and the remedy share an author.
Plausible mechanism, unverified numbers
Two forces pull opposite ways. The mechanism — behaviour degrading while correctness holds — is specific, testable, and matches what heavy users report, and the arithmetic across the reported figures is internally consistent. But it all descends from one crypto-markets retelling of an unreviewed preprint, with half the failure taxonomy unquantified and no independent run. Solid enough to go and look at your own non-thinking traffic this week; not solid enough to cite 48.64 as a constant.