Build1 distinct publisher3 min readPublished
The four-stage loop generates its own ground truth by backtesting each candidate detection against the red agent's captured telemetry, which means the fidelity of the cloned environment sets the ceiling on what it can find.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Start with the validation harness, because it is the only component in this loop with the authority to say no. The blue-agent harness generates candidate detections, and the validation harness checks each one, backtests it against the telemetry captured during the attack, returns failures for correction, and sends the survivors to the detection engine [8]. That is a compile-and-test cycle for detection content, and the test fixture is the red agent's own run. The fixture arrives labeled by construction: the red harness recorded each step of the path it executed while Falcon endpoint sensors recorded what the sensor saw [6]. The blue harness then works out which parts of the sequence it can reconstruct, which existing detections fired, and where visibility gaps remain [7]. On the defensive side those models are Nemotron builds customized for cybersecurity, running inside CrowdStrike SafeMind [2].
The stop condition is where I would push back at review. The red harness keeps adapting until it finds no further viable route inside the representative environment, and the loop ends there [10]. That environment was built from a sanitized natural-language specification of NVIDIA's accelerated computing infrastructure, which an agent-assisted workflow translated into an isolated instrumented target [11]. "Sanitized" is carrying a lot of weight in that sentence. Whatever the specification left out is a path the red agent cannot walk, and therefore a detection the blue agent is never asked to write, and the loop will still report that it ran out of routes.
The cost side of CrowdStrike's own figures is the part worth doing arithmetic on. Its internal evaluations put the Blue Solano defensive model at 97% lower cost than the leading proprietary frontier model it tested [4]. Paying three percent of the price is a ratio of about one to thirty-three [5]. For a design that retests after every deployed detection, that ratio governs how deep the loop can go before someone notices the bill. The accuracy figure is harder to use: the comparator is identified only as the leading proprietary frontier model tested, no benchmark is cited, and the metric is not defined, so whether 13% means points or a relative gain is unresolved [16].
For either number to transfer, your telemetry would have to resemble Falcon telemetry closely enough that a backtest against it means something, your detection language would have to be the one the fine-tuned Nemotron 3 Super was trained to emit [3], and your accuracy measure would have to be whatever CrowdStrike measured. The material supplied stops at the description of the instrumented target, before the per-detection results across independently seeded runs that the post says it will cover [15]. The figure I want from those runs is the false positive rate against normal enterprise activity, which NVIDIA itself names as the validation requirement [13].
So the adoption cost lands in two places. One is a rebuilt, instrumented replica of your estate. The other is a harness that can backtest a proposed rule against recorded telemetry and hand back a verdict. The manual version of this exercise is a chain of handoffs, red executes, blue reviews, engineers write, red retests, and each handoff caps how many attack variations a team gets to try [12]. If you build only the backtest, against recorded and labeled telemetry you already hold, you recover most of the iteration speed without owning a clone. That part does not require agents.
Ranked by verification strength, evidence, and original report placement.
NVIDIA and CrowdStrike evaluated an agentic attack-defense system in an isolated environment modeled on NVIDIA accelerated computing infrastructure.
On the defensive side, NVIDIA Nemotron models customized for cybersecurity operate within CrowdStrike SafeMind, CrowdStrike's agentic cybersecurity system.
For this evaluation, the optimized open-model configuration paired NVIDIA Nemotron 3 Ultra for defensive orchestration with a fine-tuned Nemotron 3 Super for detection generation.
In the first stage, starting from a threat-informed objective, the red-agent harness selected and executed an attack path inside the representative environment; its action trace recorded each step while CrowdStrike Falcon endpoint sensors captured the corresponding telemetry.
In the second stage, the blue-agent harness received the action trace, sensor telemetry and broader attack context, and using available data source information plus CrowdStrike detection-engineering expertise as grounding context, determined which parts of the event sequence could be reconstructed, which existing detections triggered, and where visibility or detection gaps remained.
In the third stage, the blue-agent harness generated candidate detections; the validation harness checked each candidate, backtested it against the captured telemetry, returned failures for correction, and sent validated detections to the detection engine.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
1.5% of Hugging Face repos take 99.2% of downloads, and the ceiling is Chinese1 distinct publisher
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
invest
Two judges, 42 hours: Nvidia's print and Warsh's first keynote price the same trade1 distinct publisher
build
A 30B model with 3B active arrives on JumpStart, aimed at the cheap middle of agent work1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Rich on method, silent on outcome
NVIDIA's post is generous about mechanism and stingy about results. The four stages, the model split, the schema knowledge base and the linting checks are all described in enough detail to be reimplemented — and then the text runs out before a single detection's performance across independently seeded runs is shown, despite promising exactly that. What remains quantitative is one sentence of CrowdStrike's own numbers against a competitor it declines to name.
One lab, one estate, no customers
Everything observed so far happens in one place: an isolated replica of NVIDIA's own infrastructure, built from NVIDIA's own sanitized spec, reviewed by NVIDIA's own security experts. The one durable adoption fact is that cybersecurity-tuned Nemotron models are wired into CrowdStrike SafeMind. No customer, no pilot, no availability date, and no second environment appears anywhere in this reporting.
Numbers outrun the proof
Overstatement here is not in the prose, which is careful and even concedes that the end-to-end cycle still takes real manual effort. It is in the arithmetic: a thirty-fold cost advantage and a 13% accuracy edge over an unnamed frontier model, published where the supporting results are not. Add the framing of a loop that runs 'until no viable path remains' — true only of a modeled environment whose fidelity nobody outside NVIDIA can inspect — and the claims sit some way ahead of what has been shown.
Both parties win if you believe it
The publisher sells the models that come out ahead, the partner sells the platform they run inside, and the loser of the comparison is a proprietary competitor left conveniently anonymous. Even the test environment is the publisher's own infrastructure, described by the publisher, vetted by the publisher's staff. None of that makes the work wrong; it does mean no step in the chain was performed by anyone with a reason to find a negative result.
Sure what was said, unsure what it means
We can be firm about provenance and design: one detailed vendor account, unambiguous about who built what and how the loop turns. We can be firm that the validation step now sits inside the attack loop and that self-generated telemetry is its ground truth. Beyond that the floor drops away — no results, no comparator, no outside reading — so any judgement about whether this actually finds gaps a human team would miss stays provisional.