Skip to content

Build1 publisher3 min readPublished

A harness gain is not a leaderboard win: reading the J-Space DeepSeek report properly

A viral X post said an inference-time text layer put DeepSeek V4 Pro ahead of Fable 5 on every task. The report it points to shows single runs, nine benchmarks, and two losses.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying A harness gain is not a leaderboard win: reading the J-Space DeepSeek report properly
Photo: runtimewire.com

What happened

  • Open-source developer Tiger380 published a benchmark report claiming a text-based control layer materially improves DeepSeek V4 Pro's performance on reasoning and agent tasks.
  • The findings drew attention on August 17th after Jun Song (@jun_song) wrote in a thread on X that the harness "completely outperforms Fable across every task."
  • The underlying report's results show DeepSeek V4-Pro-0813 gaining on all nine tested benchmarks when paired with J-Space, but do not show it beating Anthropic's Fable 5 on every task.
  • Tiger380's J-Space Cognition Suite V3.6 does not alter model weights or fine-tune DeepSeek; it packages instructions, task-routing rules and an optional Python state controller into an inference-time layer designed to keep goals and constraints active during long jobs.
  • The suite routes work into fast, full and loop modes, and for longer tasks maintains a ledger covering the goal, active facts, verified findings, open questions and next action.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Open-source developer Tiger380 published a benchmark report claiming that a text-based control layer materially improves DeepSeek V4 Pro on reasoning and agent tasks [1]. It spread on August 17th after Jun Song (@jun_song) wrote on X that the harness "completely outperforms Fable across every task" [2], which is not what the report says.

The report shows DeepSeek V4-Pro-0813 gaining on all nine tested benchmarks when paired with J-Space, but it does not show the model beating Anthropic's Fable 5 on every task [3]. Tiger380's own table puts Fable 5 ahead on Humanity's Last Exam without tools, 53.3 against 48.0 [20], a 5.3-point gap [1], and GLM-5.3 ahead on AutomationBench, 48.2 against 38.2 [21]. The report itself describes the J-Space configuration as leading the reference columns on seven benchmarks [23], which on a nine-benchmark set means it trails on two [2]. No Fable result is even listed for NL2Repo [22].

Strip out the ranking talk and the same-model half of the study is the useful part. The comparison runs the official DeepSeek Harness in its minimal configuration against the same setup plus J-Space [8]. Tool-enabled Humanity's Last Exam went from 60.0 to 67.7 [9], NL2Repo from 61.5 to 73.4 [11], DeepSWE from 62.7 to 72.0 [12], Toolathlon-Verified from 74.1 to 79.5 [14], CyberGym from 83.3 to 86.8 [15], Terminal Bench 2.1 from 87.9 to 90.1 [10], and Humanity's Last Exam without tools from 42.7 to 48.0 [16]. Agents' Last Exam and AutomationBench moved less, but positively [17]. The NL2Repo delta alone is 11.9 points [3]. According to Tiger380, base model, task, tool conditions and scoring rules were held constant, with the operating protocol as the experimental variable [18].

What is being varied is not trivial. J-Space Cognition Suite V3.6 does not alter model weights or fine-tune anything; it packages instructions, task-routing rules and an optional Python state controller into an inference-time layer meant to keep goals and constraints active during long jobs [4]. It routes work into fast, full and loop modes and maintains a ledger of the goal, active facts, verified findings, open questions and next action [5]. It also tells the model to checkpoint progress, carry diagnoses into retries and check how much of a task a test actually covers, across nine selectively loaded protocol modules plus the controller [6]. That is a structured agent protocol, not a prompt.

The uncertainty is in the statistics. Each number is one run, with no confidence intervals [7], and the report notes DeepSeek's API may ignore submitted temperature and top_p values in thinking mode [19]. Small deltas like Terminal Bench's could be run-to-run noise. The cross-model column is weaker still: comparator scores keep each vendor's own published evaluation method, and Fable, GLM, Kimi and Opus were not rerun inside one standardized J-Space evaluation [24].

One more thing worth flagging. The project name comes from Anthropic's July 6th research on internal model representations that can be reported, held and used in reasoning, studied with a technique called the Jacobian lens [25]. Tiger380's suite does not inspect or edit those activations [26]; it writes text and keeps state outside the model.

Watch for a repeated-trial version with variance bars, and for anyone rerunning Fable 5 and GLM-5.3 under the same harness and scoring rules. Until then the honest headline is that a protocol layer moved one model's reported scores, on one pass each.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories