Skip to content

Product1 publisher3 min readPublished

The cheapest model scored 10 out of 100: assistant choice is now a code-security decision

A Secure Code Warrior and RMIT study of six frontier models across 11 frameworks found no universal winner and no link between token cost and secure output.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying The cheapest model scored 10 out of 100: assistant choice is now a code-security decision
Photo: endorlabs.com

What happened

  • Secure Code Warrior and the Royal Melbourne Institute of Technology (RMIT) conducted a study evaluating the security behavior of leading AI coding models.
  • The initial study tested six frontier large language models, producing 660 complete application codebases across eleven language/framework combinations.
  • Each codebase was scanned by three independent, open-source static application security testing (SAST) tools, triaged by an AI-powered agentic false-positive verification pipeline, and scored using a multi-factor CWE Risk Model to normalize results across frameworks.
  • No model gained a universal advantage in security when measured against OWASP's Top Ten vulnerabilities, and results varied widely depending on the frameworks used.
  • GPT 5 mini, described as an efficient tool with the lowest per-token costs among the models tested, brought up the rear, scoring below average in all categories.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

Secure Code Warrior and the Royal Melbourne Institute of Technology tested six frontier large language models by generating 660 complete application codebases across eleven language and framework combinations, scanning each with three independent open source SAST tools, triaging findings through an AI false-positive verification pipeline, and scoring results with a multi-factor CWE risk model to normalise across frameworks [1][2][3]. The result that matters to anyone about to standardise one assistant across an engineering org: no model held a universal security advantage against the OWASP Top Ten, and the model with the lowest per-token cost finished last [4][5].

The overall security scores were GPT 5.1 at 79.6, Gemini 2.5 Pro at 73.5, Claude Sonnet 4.5 at 71.2, Haiku 4.5 at 49.5, Gemini 2.5 Flash at 36.4, and GPT 5 mini at 10.0 [6]. Those six are Anthropic's Sonnet 4.5 and Haiku 4.5, OpenAI's GPT 5.1 and GPT 5 mini, and Google's Gemini 2.5 Pro and Gemini 2.5 Flash [7]. The top three sit within 8.4 points of each other, and then there is a 21.7-point drop to Haiku 4.5 [1]. GPT 5.1 scored roughly eight times GPT 5 mini, and both come from the same vendor's lineup [2]. On identical tasks, the study reports, one model in some cases produced twice as much secure code as others [8].

The framework splits are where procurement gets awkward. Sonnet 4.5 led in C# Basic, Java EE JSP, Java Spring, Java Spring API and JavaScript React, while Gemini 2.5 Pro led in Python Basic, Python Django and the Swift iOS SDK [9][10]. The published text available to us breaks off mid-sentence while listing where GPT 5.1 dominated, so its framework strengths are not fully readable in this summary [11]. The category profiles differ in shape too: GPT 5.1 was above average in every OWASP category, strongest on software integrity and on security logging and monitoring failures, weakest on insecure design [12]. Gemini 2.5 Pro was the inverse, topping insecure design and identification and authentication failures while falling short on integrity failures and logging and monitoring [13]. Sonnet 4.5 had the most balanced profile with no dramatic strengths or weaknesses [14]. Gemini 2.5 Flash was below average almost everywhere but elevated on server-side request forgery, and GPT 5 mini was below average in all categories [15][16].

On cost, the study's line is blunt: models that consume more tokens or make more tool calls can cost dramatically more without producing proportionally safer code, and there is no correlation between cost and security [5][17]. That cuts both ways. It undercuts the assumption that the premium tier buys defensible code, and it undercuts the cheaper assumption that a mini model is simply a slower path to the same output.

Two caveats worth holding. The study is written in the first person by Secure Code Warrior, a vendor whose business is secure coding, so the framing favours evaluation and developer skill as the answer [18][19]. And three SAST tools across 660 codebases implies 1,980 scans, with roughly ten codebases per model-framework pair, which is thin ground for ranking within a single framework [3][4].

Watch whether the full study publishes per-framework score tables and the actual token and tool-call costs, whether anyone outside the two authors reproduces the ordering, and whether the cheap tiers that scored worst are the same ones teams are wiring into unattended agentic loops.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories