Skip to content

Build2 publishers3 min readPublished

Ornith-1.5 moves the RL loop upstream, and the hard job becomes reward design

The model now writes its own training tasks and grading harnesses. That removes the bottleneck of hand-built tasks and replaces it with a harder one: rewards that cannot be gamed.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Ornith AI released Ornith-1.5 on August 19, 2026, pushing its open model family beyond generating answers and scaffolds into generating the training tasks themselves.
  • Ornith-1.5 extends the self-scaffolding framework from Ornith-1.0 into a closed self-improvement loop: it proposes the tasks itself, generates a task-specific scaffold for each, and produces the solution rollouts used for reinforcement learning.
  • Ornith AI is betting that models can expand their own training curriculum, reducing dependence on hand-built tasks while making reward design even more consequential.
  • Ornith-1.0, released in June, learned to generate the scaffold around a coding task: the instructions, tools, decomposition and orchestration that guide a model through a longer job.
  • Each training cycle starts with an environment or codebase, broad instructions about the desired type of problem, and a history of what the model has already solved.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Ornith AI released Ornith-1.5 on August 19, 2026, extending the self-scaffolding framework from Ornith-1.0 into a closed loop in which the model proposes its own training tasks, generates the scaffold and evaluation harness for each one, and produces the rollouts it then learns from [1][2]. The consequence is a shift in where the scarce engineering effort sits: Ornith AI's own framing is that this reduces dependence on hand-built tasks while making reward design even more consequential [28]. Ornith-1.0, shipped in June, learned to write the scaffold around a coding task, meaning the instructions, tools, decomposition and orchestration that carry a model through a longer job [3]. Ornith-1.5 adds the assignment itself. Each training cycle begins with an environment or codebase, high-level instructions about the type of problem wanted, and the model's own history of solved problems [4]. The system proposes a harder task, generates or refines the scaffold, and produces a rollout conditioned on both, with reward propagating back across all three stages so that question, harness and answer improve together [2][5]. The scoring is where the design shows its hand. Task reward multiplies three signals: whether task and scaffold form a valid and verifiable environment, whether difficulty sits near the current capability frontier, and whether the task is novel against work already generated [6]. Frontier difficulty targets a 0.2 empirical success rate, so a task loses value to the generator once the model clears it reliably [7]. That implies roughly four rollouts in five are expected to fail [19], which is the point: hard enough to teach, with enough successes left to train on [8]. Validity acts as a hard gate that zeroes out malformed tasks, and all three stages are optimised with GRPO, according to Ornith [9]. The harness gets its own reward for matching the assignment, measuring solution quality accurately, and resisting reward hacking [10]. This is the load-bearing part. A system that writes both the exam and the grading script can manufacture easy points through malformed tests, hidden shortcuts or criteria that reward the wrong behaviour [11]. It is also the part with the least public detail: the official Ornith-1.0 repository documents benchmark harnesses, anti-hacking filters and resource settings, but runtimewire reports no substantiation of a more specific evaluation protocol involving measures such as Git-history removal or network blocking [12]. One word deserves deflating. Self-improvement here describes the training procedure, not the shipped artefact: the downloaded model does not continuously retrain itself on a user's phone or workstation, and Ornith runs the loop during development before publishing weights [13]. The release covers a 397B mixture-of-experts flagship, a 35B MoE model activating 3B parameters per token, and a 9B dense model with a quantized mobile build for iPhone and Android [14]. On the company's published tables, averaged over five independent runs, the 397B scores 85.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, which Ornith reports as on par with Claude Opus 4.8 at 85.0 and 59.0 [15]: ahead by 0.1 on one, behind by 3.0 on the other [18]. The 35B reaches 68.5 on Terminal-Bench 2.1 and 79.0 on SWE-Bench Verified, the 9B 47.0 and 70.6 [16], and the flagship posts 92.8 on GPQA Diamond and 86.6 on BrowseComp [17]. Ornith AI documents local deployment for the 9B and publishes an MLX build for Apple hardware [25]. The lineage is consistent with the bet. Ornith-1.0 shipped in June 2026 in four variants under an MIT license with weights on Hugging Face, post-trained on Gemma 4 and Qwen 3.5 checkpoints, and introduced the scaffold as a learnable object co-evolving with the policy [20][21]. DeepReinforce has published reinforcement learning optimisation work openly before, including CUDA-L1 and the IterX agent loop, with reward hacking a running concern [22]; runtimewire describes a related kernel-optimisation effort that used execution speed as the reward but names it CUDA-L2 [23], a discrepancy to keep in mind when tracing the research trail [27].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories