Product1 distinct publisher3 min readUpdated
A Secure Code Warrior and RMIT study of six frontier models across 11 frameworks found no universal winner and no link between token cost and secure output.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
Secure Code Warrior and the Royal Melbourne Institute of Technology tested six frontier large language models by generating 660 complete application codebases across eleven language and framework combinations, scanning each with three independent open source SAST tools, triaging findings through an AI false-positive verification pipeline, and scoring results with a multi-factor CWE risk model to normalise across frameworks [1][2][3]. The result that matters to anyone about to standardise one assistant across an engineering org: no model held a universal security advantage against the OWASP Top Ten, and the model with the lowest per-token cost finished last [4][5].
The overall security scores were GPT 5.1 at 79.6, Gemini 2.5 Pro at 73.5, Claude Sonnet 4.5 at 71.2, Haiku 4.5 at 49.5, Gemini 2.5 Flash at 36.4, and GPT 5 mini at 10.0 [6]. Those six are Anthropic's Sonnet 4.5 and Haiku 4.5, OpenAI's GPT 5.1 and GPT 5 mini, and Google's Gemini 2.5 Pro and Gemini 2.5 Flash [7]. The top three sit within 8.4 points of each other, and then there is a 21.7-point drop to Haiku 4.5 [1]. GPT 5.1 scored roughly eight times GPT 5 mini, and both come from the same vendor's lineup [2]. On identical tasks, the study reports, one model in some cases produced twice as much secure code as others [8].
The framework splits are where procurement gets awkward. Sonnet 4.5 led in C# Basic, Java EE JSP, Java Spring, Java Spring API and JavaScript React, while Gemini 2.5 Pro led in Python Basic, Python Django and the Swift iOS SDK [9][10]. The published text available to us breaks off mid-sentence while listing where GPT 5.1 dominated, so its framework strengths are not fully readable in this summary [11]. The category profiles differ in shape too: GPT 5.1 was above average in every OWASP category, strongest on software integrity and on security logging and monitoring failures, weakest on insecure design [12]. Gemini 2.5 Pro was the inverse, topping insecure design and identification and authentication failures while falling short on integrity failures and logging and monitoring [13]. Sonnet 4.5 had the most balanced profile with no dramatic strengths or weaknesses [14]. Gemini 2.5 Flash was below average almost everywhere but elevated on server-side request forgery, and GPT 5 mini was below average in all categories [15][16].
On cost, the study's line is blunt: models that consume more tokens or make more tool calls can cost dramatically more without producing proportionally safer code, and there is no correlation between cost and security [5][17]. That cuts both ways. It undercuts the assumption that the premium tier buys defensible code, and it undercuts the cheaper assumption that a mini model is simply a slower path to the same output.
Two caveats worth holding. The study is written in the first person by Secure Code Warrior, a vendor whose business is secure coding, so the framing favours evaluation and developer skill as the answer [18][19]. And three SAST tools across 660 codebases implies 1,980 scans, with roughly ten codebases per model-framework pair, which is thin ground for ranking within a single framework [3][4].
Watch whether the full study publishes per-framework score tables and the actual token and tool-call costs, whether anyone outside the two authors reproduces the ordering, and whether the cheap tiers that scored worst are the same ones teams are wiring into unattended agentic loops.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Secure Code Warrior and the Royal Melbourne Institute of Technology (RMIT) conducted a study evaluating the security behavior of leading AI coding models.
The initial study tested six frontier large language models, producing 660 complete application codebases across eleven language/framework combinations.
Each codebase was scanned by three independent, open-source static application security testing (SAST) tools, triaged by an AI-powered agentic false-positive verification pipeline, and scored using a multi-factor CWE Risk Model to normalize results across frameworks.
No model gained a universal advantage in security when measured against OWASP's Top Ten vulnerabilities, and results varied widely depending on the frameworks used.
GPT 5 mini, described as an efficient tool with the lowest per-token costs among the models tested, brought up the rear, scoring below average in all categories.
Overall security scores: GPT 5.1 79.6; Gemini 2.5 Pro 73.5; Sonnet 4.5 71.2; Haiku 4.5 49.5; Gemini 2.5 Flash 36.4; GPT 5 mini 10.0.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Structured methodology, unverifiable specifics
The study describes a real and reasonably rigorous design - 660 generated codebases, eleven framework combinations, three independent SAST tools, agentic false-positive triage, CWE-normalized scoring - which is more than most model-comparison claims carry. But every number reaches the reader through a single self-authored article with no link to the study, no task set, no raw scans, no variance or repetition counts, and only about ten codebases per model-framework cell. Qualitative category findings ('above average', 'elevated SSRF') have no supporting figures, and the cost half of the argument reports no per-model token or dollar data at all.
No adoption signal in sources
The only observable event is publication of a benchmark result. The supplied source reports no deployments, no team or organization changing model selection on the basis of these findings, no usage disclosures, and no uptake of the scoring model by anyone outside the authors. Adoption cannot be measured without inferring facts the source does not contain.
Mildly overstated certainty
The article's own framing is restrained - it argues there is no winner rather than crowning one - and its headline conclusions (framework-dependence, price does not equal security) are the sort of finding that resists overselling. The gap comes from precision theatre: single-decimal scores and a dramatic 10.0-out-of-100 result presented as settled measurement, with no data release, no variance, no replication and no vendor rebuttal, all authored by a company that sells secure-coding services. Directionally plausible, more precise-sounding than the disclosed evidence supports.
Vendor-authored, commercially aligned
The study is produced and written up in the first person by Secure Code Warrior, a company whose business is developer secure-coding capability, and its findings - that AI-generated code security is unpredictable, framework-dependent and unrelated to price, so teams need contextual evaluation and developer skill - map directly onto what such a vendor sells. The RMIT partnership and the publication of unflattering results for models across all three major labs, with no product pitch in the supplied text, moderate but do not remove the alignment. No disclosure of the commercial interest appears in the article.
Single self-authored source
Confidence is capped by cluster structure: one publisher, one article, authored by the study's own team, with no replication, no raw data, no per-cell sample detail and no response from the model vendors whose products are ranked. The methodology description is specific enough to be credible in outline and the derived arithmetic is internally consistent, so the directional claims can be reported with attribution - but the individual scores should not be treated as verified.
invest
Google Ships Flash Instead of Pro While OpenAI Loses Its Two Best Operators1 distinct publisher
build
Solar Pro 4 turns model routing into a procurement decision, not a research one1 distinct publisher
product
Incogni ranks 13 AI assistants by privacy risk: bigger is worse, except ChatGPT1 distinct publisher
invest
Korea's 720-billion-dollar AI plan meets its first critic: the man running its research hub1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 19, 2026