Skip to content

Build1 publisher2 min readPublished

Cantina's open-weights exploit model ran a 60-task security eval for $2.38

Cantina released apex-flash-1, an open-weights vulnerability-research model it says solved 40 of 60 tasks for $2.38, against $74.68 for Claude Opus 5 High. There is no hosted endpoint, so teams download the 321-billion-parameter weights, pay for their own inference and verify the numbers themselves.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Cantina's open-weights exploit model ran a 60-task security eval for $2.38
Generated illustration

What happened

  • apex-flash-1 is a reinforcement-learning fine-tune of the open GLM-5.3-Flash base model, and Cantina has put the weights up to download on Hugging Face.
  • On the same 60-task set, Claude Opus 5 High solved 43 and the untuned GLM-5.3-Flash solved 36, so Cantina's model placed below Claude and above the base it was trained from.
  • Cantina came out of stealth in July with an $8 million round led by Framework Ventures.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost The $2.38 covers only Cantina's own eval run; with no token-priced endpoint, whoever adopts the model pays inference on 321 billion parameters on their own hardware.
  • constraint A model this size with no hosted API rules out a quick call to try it, so testing the cost claim means standing up your own inference capacity first.
  • exposure The public weights carry a model trained to pursue and verify exploits in a running target, so defenders and attackers can download the same capability.

The $2.38 is the cost of running the test, not of using the model on real targets. Cantina ran its 60-task evaluation for that figure and reported $74.68 to run the same tasks on Claude Opus 5 High, about 31 times more [10][2]. Claude also solved more of them: 43 against apex's 40, or 71.7 percent against 66.7 [8][7][1][3].

The comparison that argues for the training is the base model. Untuned GLM-5.3-Flash solved 36 of the same tasks and cost $4.56 to run them [9][10]. Cantina's reinforcement-learning fine-tune of that model solved 40 and cost $2.38, roughly half [5][7][10][5]. It is both more capable and cheaper than the model it was trained from [5].

Every one of those numbers comes from Cantina's own model card [21]. The 60 tasks were drawn from 20 held-out vulnerability cases and run in isolated environments, with verifiers that checked the final state of each target [6][11]. Mulackal argues public cyber benchmarks are poor proxies for the flaws customers actually hit, so Cantina built its own around offensive and defensive tasks [18]. That gives the company a measure it considers closer to real work. It also means no outsider can line apex up against Claude on a neutral test [22].

There is no hosted endpoint. In a reply, Mulackal said Cantina has no hosted option billed per token for now, and Cantina points users at an agent harness such as Codex [4][3]. The download is free; the running is not. A 321-billion-parameter model is inference you provision and pay for yourself, so whether the $2.38 transfers depends on your hardware and your targets [5].

Mulackal worked on the Solidity language and compiler at the Ethereum Foundation before co-founding Cantina and the security firm Spearbit, according to his Devcon profile [13]. In his October 4 post he framed security as an economics problem and cited Anthropic's September 10 report on AI-assisted "exploit foundries" as evidence the work is being automated [1][14][15]. He also said Cantina has earned $1 million in bug bounties, ranks first on HackerOne's US business leaderboard for 2026, and that its harnesses process trillions of tokens a month [16][17]. Those are claims in the post, not audited figures. Cantina came out of stealth in July with an $8 million round led by Framework Ventures [19].

What to watch

  • Whether any independent benchmark reproduces the cost and pass-rate gap on targets Cantina did not design.
  • Whether Cantina ships hosted, token-priced access that removes the self-hosting burden.
  • Whether Cantina's bug-bounty and leaderboard claims are backed by audited, outcome-level figures.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories