Build1 distinct publisher3 min readUpdated
A viral X post said an inference-time text layer put DeepSeek V4 Pro ahead of Fable 5 on every task. The report it points to shows single runs, nine benchmarks, and two losses.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Open-source developer Tiger380 published a benchmark report claiming that a text-based control layer materially improves DeepSeek V4 Pro on reasoning and agent tasks [1]. It spread on August 17th after Jun Song (@jun_song) wrote on X that the harness "completely outperforms Fable across every task" [2], which is not what the report says.
The report shows DeepSeek V4-Pro-0813 gaining on all nine tested benchmarks when paired with J-Space, but it does not show the model beating Anthropic's Fable 5 on every task [3]. Tiger380's own table puts Fable 5 ahead on Humanity's Last Exam without tools, 53.3 against 48.0 [20], a 5.3-point gap [1], and GLM-5.3 ahead on AutomationBench, 48.2 against 38.2 [21]. The report itself describes the J-Space configuration as leading the reference columns on seven benchmarks [23], which on a nine-benchmark set means it trails on two [2]. No Fable result is even listed for NL2Repo [22].
Strip out the ranking talk and the same-model half of the study is the useful part. The comparison runs the official DeepSeek Harness in its minimal configuration against the same setup plus J-Space [8]. Tool-enabled Humanity's Last Exam went from 60.0 to 67.7 [9], NL2Repo from 61.5 to 73.4 [11], DeepSWE from 62.7 to 72.0 [12], Toolathlon-Verified from 74.1 to 79.5 [14], CyberGym from 83.3 to 86.8 [15], Terminal Bench 2.1 from 87.9 to 90.1 [10], and Humanity's Last Exam without tools from 42.7 to 48.0 [16]. Agents' Last Exam and AutomationBench moved less, but positively [17]. The NL2Repo delta alone is 11.9 points [3]. According to Tiger380, base model, task, tool conditions and scoring rules were held constant, with the operating protocol as the experimental variable [18].
What is being varied is not trivial. J-Space Cognition Suite V3.6 does not alter model weights or fine-tune anything; it packages instructions, task-routing rules and an optional Python state controller into an inference-time layer meant to keep goals and constraints active during long jobs [4]. It routes work into fast, full and loop modes and maintains a ledger of the goal, active facts, verified findings, open questions and next action [5]. It also tells the model to checkpoint progress, carry diagnoses into retries and check how much of a task a test actually covers, across nine selectively loaded protocol modules plus the controller [6]. That is a structured agent protocol, not a prompt.
The uncertainty is in the statistics. Each number is one run, with no confidence intervals [7], and the report notes DeepSeek's API may ignore submitted temperature and top_p values in thinking mode [19]. Small deltas like Terminal Bench's could be run-to-run noise. The cross-model column is weaker still: comparator scores keep each vendor's own published evaluation method, and Fable, GLM, Kimi and Opus were not rerun inside one standardized J-Space evaluation [24].
One more thing worth flagging. The project name comes from Anthropic's July 6th research on internal model representations that can be reported, held and used in reasoning, studied with a technique called the Jacobian lens [25]. Tiger380's suite does not inspect or edit those activations [26]; it writes text and keeps state outside the model.
Watch for a repeated-trial version with variance bars, and for anyone rerunning Fable 5 and GLM-5.3 under the same harness and scoring rules. Until then the honest headline is that a protocol layer moved one model's reported scores, on one pass each.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Open-source developer Tiger380 published a benchmark report claiming a text-based control layer materially improves DeepSeek V4 Pro's performance on reasoning and agent tasks.
The underlying report's results show DeepSeek V4-Pro-0813 gaining on all nine tested benchmarks when paired with J-Space, but do not show it beating Anthropic's Fable 5 on every task.
Tiger380's J-Space Cognition Suite V3.6 does not alter model weights or fine-tune DeepSeek; it packages instructions, task-routing rules and an optional Python state controller into an inference-time layer designed to keep goals and constraints active during long jobs.
The suite routes work into fast, full and loop modes, and for longer tasks maintains a ledger covering the goal, active facts, verified findings, open questions and next action.
The suite directs the model to checkpoint progress, carry diagnoses into retries and verify how much of a task a test actually covers; the repository contains nine selectively loaded protocol modules alongside the optional controller.
Each reported result represents one run rather than an average across repeated trials, and the report provides no confidence intervals.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but single-run and self-published
The cluster contains precise per-benchmark before/after numbers and a genuinely controlled same-model A/B design, which is more evidence than most harness claims carry. It is undercut by the fact that every score is one run with no confidence intervals, sampling parameters may not have been honored, the report is authored by the suite's own developer, no independent replication exists, and only one publisher covers it.
Artifact published, no usage evidence
What is observable is that the suite is publicly available as a repository and that its benchmark report circulated on X. The supplied material contains no deployment counts, downloads, stars, integrations, pricing or user disclosures, so real uptake cannot be sized beyond publication plus social attention.
Viral framing well ahead of the report
The circulating claim — that the harness 'completely outperforms Fable across every task' — is materially stronger than the report it cites, which shows leads on seven of nine reference columns, losses to Fable 5 on Humanity's Last Exam without tools and to GLM-5.3 on AutomationBench, single-run measurements, and comparator scores never rerun under one method. The gap is positive and large, though partially offset because the underlying same-model gains are real as reported and the reporting itself corrects the overstatement.
Author evaluates own tool; amplifier boosts it
The benchmark report is produced by the same developer who publishes the J-Space suite, so the favorable same-model result is self-assessment. A third party then restated it in a stronger form on X, and the comparator column reuses scores each vendor published under its own method, meaning vendor self-reporting is embedded in the table too. Offsetting this, the author documents limitations and warns against reading the table as a head-to-head test.
Single publisher, transparent limits
Confidence is limited by the cluster containing exactly one publisher and one underlying self-published report, with no replication or vendor response. It is raised somewhat by the internal consistency of the account: specific numbers, an explicit design description, and the author's own caveats about single runs, sampling parameters and non-standardized comparators.
build
GLM-5.3 kept the base model and bought ten times the environments instead2 distinct publishers
build
GLM-5.3 changed nothing but the training environments. That is the whole test.3 distinct publishers
build
Open weights caught up on finding bugs. They did not catch up on using them.1 distinct publisher
build
OpenAI's president says open weights will accelerate the threat. His own cyber model stays gated.1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 17, 2026