Skip to content

Build1 publisher2 min readPublished

Cantina publishes a security model built from 50 of its own vulnerability cases

Cantina released Apex Flash-1, an open-weights security model fine-tuned on 50 of its own vulnerability cases. Its only performance evidence is a 60-task evaluation the company ran itself, and a separate build is modified to refuse fewer requests.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Cantina publishes a security model built from 50 of its own vulnerability cases
Generated illustration

What happened

  • Cantina estimates a single evaluation run cost $2.38 on its own model, $4.56 on the base model and $74.68 on Claude Opus 5 High, using provider pricing.
  • The downloadable checkpoint is a 321-billion-parameter model under an MIT license that runs locally on common inference software.
  • Cantina raised $8 million led by Framework Ventures in July, bringing total funding to $16.5 million, and launched a security operations platform called Clarion.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure The weights download freely and are not limited to the authorized researchers Cantina names as its users, so defenders and attackers can run the same security tool.
  • constraint The comparison does not cover other codebases, security teams or agent setups, so a buyer cannot read the pass rate as performance on their own stack.
  • decision Choosing Apex Flash-1 over Claude Opus 5 High trades about five points of measured accuracy for a run roughly 31 times cheaper, a trade that holds only if the workload resembles Cantina's tasks.

The model is post-trained on Cantina's own vulnerability work, built with Yeta Labs [1][6]. The evaluation Cantina published on October 1 runs 60 tasks built on 20 cases held out of training, once per model, on environments Cantina built [3][9][14]. Apex Flash-1 passed 40, a 66.7% pass-at-one score; the base GLM-5.3-Flash it is fine-tuned from passed 36, or 60%, and Claude Opus 5 High passed 43, or 71.7% [10][11][12][4]. With 60 tasks, each pass moves the score by about 1.67 points, so the five-point gap to Opus is three solved tasks and the edge over the base model is four [4][2][1]. A different held-out set could move either margin.

By Cantina's own figures the fine-tune ran at about half the cost of the base model it came from [5]. Both are the same 321-billion-parameter architecture, so the cheaper run reflects fewer tokens per task, not a smaller model [4][5].

The training set was built to test whether an agent can do more than flag a suspicious code pattern [8]. Cantina made 150 tasks from 50 vulnerability cases, rendering each case three ways: source code with guidance, source code with little direction, and a running target with no source and little direction [7]. In one case the model held access to one Forgejo workflow and had to pull a protected artifact from another, tracing a mismatch between what a signed download URL authorized and what the application then fetched, and showing the effect on a live system [16]. A verifier checked whether the artifact was actually recovered, and passing runs were reviewed for shortcuts [17]. Cantina frames the model as a worker that a larger agent points at one such task [18].

The reported scores cover the standard model, not the abliterated build [19][2]. Cantina's argument for releasing open weights is that local deployment gives defenders more control over their data, tools and model behavior [20], and that attackers can already adapt open models while legitimate teams face restrictions on proprietary services [21]. runtimewire reports that fewer refusals may help authorized researchers fit the model to their workflow while weakening a barrier against harmful requests [22].

The release also turns Cantina's vulnerability research into a second product [28]. Mulackal founded the company after building Spearbit, a researcher network that grew from digital-asset audits into contests and bug-bounty work, and Cantina now sells autonomous security software alongside researcher-led engagements [24][25]. Its co-founder and president, Mike Leffer, spent more than a decade in defense and security operations and served in the U.S. Army [26].

What to watch

  • Whether an independent benchmark reproduces the 66.7% pass rate on codebases Cantina did not build.
  • Whether the abliterated weights turn up in real attacks, and how Cantina responds.
  • Whether Cantina releases the held-out cases or task-level results for outside review.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories