Build1 distinct publisher3 min readUpdated
Veracode ran more than 150 models over 80 tasks and found 45% of the output carries a known weakness. The level is bad; the flat two-year trend is the part that changes your plan.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The two headline numbers are one number. Forty-five percent of outputs carrying a known weakness and a pass rate of about 55% are the same measurement read from opposite ends [1]. So the quantity worth arguing about is not the level, it is the slope, and Veracode puts that at roughly zero percentage points a year across two years of model releases [5][4].
The language spread is where an actual decision sits. Java passed 29% of the time, which means about 71% of Java outputs carried a known weakness, against 38% for Python at its 62% pass rate: close to 1.9 times the defect density from the same models on the same style of task [7][2]. One caution on the arithmetic. The four published language figures average 51.5%, which is below the 55% aggregate [6], so the headline is not a flat mean of the four and the per-language task mix is the thing to ask about before anyone builds policy on the Java figure alone.
Park's mechanism for this is that the models are not reasoning about security, they are reproducing the posture of whatever they saw most of: SQL injection has been the teaching example since roughly 2005 and parameterised queries are everywhere in the corpus, while log injection is a real weakness that never became a lesson [11]. If that is right, the lever is corpus composition, which nobody downstream of a lab controls. It also fits what the report found at the top end. The reasoning-tuned models did better at 70 to 72% [6], which still leaves close to three outputs in ten carrying a known weakness [3].
That pushes the load onto review, and review is where the Stanford result bites. Neil Perry and colleagues found participants with an AI assistant, codex-davinci-002 in that study, wrote significantly less secure code [8], and were more likely to believe they had written secure code [9]. The same study found the mitigation: participants who distrusted the assistant and rewrote their prompts produced fewer vulnerabilities [10]. Skepticism is a personal habit, not a control you can install across a team on a deadline.
The tooling argument follows from timing rather than accuracy. SAST reads source at rest, DAST hits a running application, IAST instruments runtime, and all three assume the code already exists [13]. In a semi-autonomous agent turn, a call like subprocess.run(user_input, shell=True) can be written, saved and executed before any gate runs, with the pre-commit hook firing afterwards if it fires at all [14].
Two things to hold at arm's length. Park is the creator of Cencurity, an open-source security gateway for LLM coding agents [15], so the conclusion that control belongs inside the loop is also his product thesis. And his four-surface framing, prompt injection, unsafe tool calls, outbound data leakage, and actions with no record, is his enumeration rather than Veracode's measurement [12]. The claim that code scanning addresses one surface in four [5] is a coverage argument, not a measured 25%. The 45% is measured [4]; the other 75% of the problem is asserted.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Veracode published its Spring GenAI Code Security Update in March 2026.
Veracode ran more than 150 large language models through 80 code-generation tasks in Java, JavaScript, C# and Python, then tested every output against four common weakness categories.
Over 95% of the generated code compiled and ran.
Forty-five percent of the generated code contained a known vulnerability.
Veracode states that security pass rates 'remain stubbornly stuck at approximately 55%', virtually identical to where they stood two years earlier.
Reasoning-tuned models performed better, reaching 70 to 72% security pass rates.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Named third-party studies, relayed by one interested author
The core quantitative claims are attributed to identifiable research — Veracode's March 2026 Spring GenAI Code Security Update and a Stanford controlled user study led by Neil Perry — with specific methodology (150+ models, 80 tasks, four languages) and consistent internal arithmetic, which is well above anecdote. But the cluster contains exactly one item, no primary link or citation to either study, no per-weakness-class figures despite the argument resting on them, an unexplained gap between the four per-language rates (mean 51.5%) and the 55% aggregate, and no independent verification of the author's own taxonomy, one-turn execution scenario or gateway behaviour.
No deployment or usage evidence supplied
The only observation in the cluster is a third-party benchmark of model output quality, which measures models, not uptake. Nothing is supplied about installs, stars, users, enterprise deployments, downloads or revenue for Cencurity or the 'CAST' category, and no adoption figures are given for the agent tools mentioned (Roo Code, Claude Code) or for the practice of gateway-based agent security. Adoption cannot be scored without inferring facts the source does not provide.
Grounded diagnosis, over-claimed remedy
The diagnosis is close to aligned with its evidence: the flat two-year pass rate and the 45% vulnerable-output figure are attributed to a named study and the headline framing does not exceed them. The overstatement sits downstream. Two views of one measurement (45% vulnerable, 55% passing) are presented as separate alarms; a self-coined category ('CAST') and an undisclosed-adoption gateway are positioned as the answer with no detection, false-positive or overhead data; and the author's four-mode taxonomy is used to argue existing tooling covers only a quarter of the surface without any independent weighting of those modes.
Vendor-authored, disclosed product stake, vendor-sourced data
The piece is written by the creator of Cencurity, the product it concludes you need, and the conflict is disclosed openly in the byline and body. It also coins a category name ('CAST') that positions that product as the definitional example. The primary data source is itself an application security vendor whose commercial interest aligns with a finding that AI-generated code is unsafe, and the article does not weigh that. Disclosure is credited, but the alignment between the argument and the author's and data source's commercial interests is strong and one-directional.
Moderate-low: plausible, single-source, unverified
Confidence is limited by structure rather than plausibility. One publisher, one interested author, two studies quoted without primary links, missing per-weakness data, an unreconciled aggregate-versus-per-language arithmetic gap, and zero adoption evidence for the proposed remedy. The directional finding — that AI code security has not tracked capability gains, and that agent loops create a control window earlier than SAST/DAST/IAST operate — is internally coherent and consistent with the named studies, which keeps confidence near the middle rather than low.
build
Best agent in Databricks' live document-reasoning contest scored 63.3%2 distinct publishers
product
A 2x LLM bill is not a bug report: token spend is an observability problem1 distinct publisher
build
Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill1 distinct publisher
build
The $559M-versus-$12.3B quarter matters more than the $65B run rate4 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 21, 2026