Skip to content

Build1 publisher2 min readPublished

Four passing runs out of 80 separate first from second on Specific's private-code benchmark

Specific Labs scores coding agents on licensed production codebases. The best setup clears 38.8%. The analysis covers ten tasks at eight runs each, so every published score is a count of passing rollouts out of 80.

The Engineer · Build desk

Photograph accompanying Four passing runs out of 80 separate first from second on Specific's private-code benchmark
Photo: every.io

What happened

  • Specific Labs co-founders Janak Sunil and Siddhant Paliwal launched Real-SWE, a benchmark that sets coding agents to work on licensed, private production codebases.
  • One sample assignment asks an agent to repair invoice taxation across businesses with customer exemptions, buyer destinations, European VAT registrations and separate sandbox and production tax services.
  • Fable 5.1 through Claude Code tops the published leaderboard with a 38.8% resolution rate, and the last-placed setup, GPT-5.6 Sol with Codex CLI, scored 16.2%.
  • The published analysis covers ten tasks, eight model-and-harness combinations and eight attempts per pairing, for 640 rollouts, with resolution rate defined as pass@1 averaged over the eight runs.
  • Specific scores each model in its own native coding harness, so every entry measures a model-and-harness combination.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint One task flipping from failed-on-all-eight to passed-on-all-eight moves a combination by ten points, so the order below the top entry is only as stable as the choice of ten tasks.
  • exposure The work being failed is invoice tax and settled-sales reporting, and according to runtimewire.com a patch that looks plausible can still apply the wrong rule or touch code the running service never calls.
  • cost Specific did not publish the full task set or enough methodology to reproduce the run, so anyone who wants an independent check has to license private production codebases and build the harness themselves.
  • precedent The company publishing the failure rate also licenses operational data and codebases. Private evaluations owned by the company licensing out the data and code are a workable business shape for others to copy.

Every number on Specific's leaderboard lands on a multiple of 1.25 points. Ten tasks at eight runs apiece give each combination 80 rollouts, and one rollout is worth 1.25 points [20]. Fable 5.1's leading 38.8% is 31 passes out of 80, and GPT-6 Astra's 33.8% is 27 [20][21]. Four passing runs separate first from second [21]. The whole table, first to eighth, spans 22.5 points, which is 18 rollouts [23].

Six of the ten analyzed tasks scored under 15% aggregate [9]. Each of those tasks was attempted 64 times across the eight combinations, so under 15% means nine passes or fewer [26]. Two entries ran on the same harness, Codex CLI, and finished 17.6 points apart [24]. Eight independent runs per task is careful measurement, and ten tasks is still a pilot.

The reading offered by runtimewire.com is that public coding scores can reward familiarity with visible repositories, and that Real-SWE gives Specific a route from evaluation failures to training-data customers [12]. There is a reason behind the first half: public benchmarks can expose the repository, the issue and the eventual patch, any of which can enter training data, while a private codebase makes an agent read the system in front of it and infer conventions from the workspace [11]. Confirming it needs a matched pair: the same eight setups, the same native harnesses, the same pass@1 over eight runs, scored on a public benchmark next to Real-SWE. Specific has not run that comparison; the published leaderboard reports Real-SWE [3].

Across eight repository-backed sample tasks, the median instruction runs 1,742 characters and the median reference solution touches 11 files [18]. That works out to about 158 characters of instruction per file the reference changes [25]. The agent has to find the rest of the work itself [27].

Y Combinator's company page says Specific acquires and licenses operational data and codebases [13], and YC lists it as a two-person company founded in 2025 and backed in the Fall 2025 batch [14]. The sample codebases are described without being named: an events product with more than 200,000 users, a consumer fintech that has processed more than 100,000 bank statements, and enterprise sales software, all on Specific's account [15].

What to watch

  • Whether Specific publishes the task set or a reproducible harness spec so someone outside the company can recount the 640 rollouts.
  • Whether the leaderboard grows past ten tasks, since each task added cuts the ten-point weight a single flipped task now carries.
  • Whether Specific ever scores the same eight model-and-harness pairs on a public benchmark for a side-by-side comparison.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories