Build1 publisher3 min readPublished
H Company's Holo4 merges its desktop and API specialists into one open-weight model
H Company's open-weight Holo4 27B handles GUIs, code, MCP servers and REST APIs in one model, scoring 61.7% on OSWorld 2.0 at an estimated $1.22 per task. The choice between specialists happens during training, so agent teams could drop the router from their stack if the scores hold on their own workloads.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- H Company released Holo4 on September 28, 2026, in two sizes: a 27B dense model and a 35B-A3B mixture-of-experts model.
- Its training tasks came from an Agentic Task Factory that built about 10,000 verifiable tasks from product manuals, help articles, website screenshots and open-source software.
- Supervised fine-tuning ran on 127B tokens, roughly three quarters of them successful agent trajectories produced by that factory.
- Closed models still lead OSWorld 2.0, with Claude Opus 5.5 at 81.8% and GPT-6 Astra at 73.5%, both at significantly higher inference cost according to the writeup.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Picking the 35B MoE for high-volume desktop work saves money only if failed runs cost less than finished ones, because per completed task the two sizes cost about the same.
- constraint Agents that mostly call MCP servers and REST APIs depend on the smallest slice of the training mix, so the one-model claim is least proven on the work a separate tool-calling model would do.
- decision Teams weighing Holo4 against a specialist-plus-router stack have to run that comparison themselves, since H Company's reported wins are against its own earlier models.
Holo4 still has specialists inside it. In the reinforcement learning stage, H Company trained one LoRA expert on desktop and web environments and a second on terminal, MCP and API environments [16]. According to the dev.to writeup, keeping them apart lets each expert build strong priors for its domain without the two interfering during RL [16]. Both experts were then merged back into the fine-tuned base model with equal weight [17]. At each step, the merged model chooses between screenshot clicks, sandboxed code, MCP calls and direct API requests from one shared action space [9].
The writeup says most labs handle mixed interfaces by building separate models or by routing between specialists at inference time [8]. A router is one more model to evaluate in production. A misrouted task also fails before the right specialist ever sees it. Holo4 makes that choice once, in training, and that leaves one fewer model to get paged about. The cost is a fixed blend. A team that only needs the API side still serves weights that carry both experts [17].
The training mix favors GUI work. The model card puts desktop at 45% of training tokens, web at 14%, MCP/API at 12%, mobile at 3% and other at 26% [18]. Desktop and web together come to 59%, nearly five times the MCP/API share [20]. The task factory's own output is more even, at 40% web app tasks, 30% MCP server tasks and 30% desktop and OS tasks [12].
Both sizes use Qwen3.8 as the base [10], so the multi-interface behavior comes from data and training. The factory's quality gate is the best engineering in the release. A task stays only if its verifier fails on the untouched seed state, passes on the golden state, rejects every near-miss, and an agent can solve it through the real interface [13]. The near-miss rule tests the verifier itself. A task cannot be passed by a check that accepts almost-right states. Failed attempts go back into an audit and hardening loop instead of being thrown away [14]. The writeup contrasts this with the common practice of filtering trajectories by outcome alone [14].
The benchmark figures are H Company's claims about its own evaluation setup. The company reports that Holo4 beats its previous models on every benchmark it tested, including 85.2% on the original OSWorld, 85.1% on AndroidWorld and 45.4% on AutomationBench [1]. OSWorld 2.0, the source of the lower 27B score, covers long multi-step desktop workflows [2]. The writeup does not include a comparison against a GUI specialist and a tool-calling model behind a router, which is the setup Holo4 is pitched against. For the OSWorld 2.0 result to transfer, your tasks need to resemble its desktop workflows. For the cost figure to transfer, your serving cost per task needs to match whatever produced H Company's estimate [2].
The writeup calls the 35B-A3B MoE the more cost-efficient option for high-volume deployments [5]. It scores 30.9% on OSWorld 2.0 at $0.61 per task [4]. Divide each model's cost per task by its success rate and the 27B comes to about $1.98 per completed task, the MoE to about $1.97 [19]. Per attempt, the MoE is half the price. Per finished task, the two cost the same, provided a failed run costs as much as a successful one and a retry succeeds at the benchmark rate [19].
What to watch
- Independent runs of Holo4 27B on OSWorld 2.0 with published hardware and per-token costs, to test the $1.22 per-task estimate.
- A same-task comparison of Holo4 against a GUI specialist and a tool-calling model behind a router.
- Whether H Company publishes the two LoRA experts separately, so API-only teams could serve the terminal, MCP and API expert alone.