Skip to content

Build1 publisher3 min readPublished

CrowdStrike's own triage numbers make AI auto-close a calibration contract, not a headcount cut

High-confidence "no problem" precision fell from a roughly 98% target to 90.8% one week past tuning. Closing alerts automatically buys you a monitoring job, not fewer analysts.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • CrowdStrike published research titled "Teaching AI to Reason Through Detection Triage" with a publication date of 2026-08-17, on having AI judge whether Windows endpoint alerts are real attacks or harmless false positives.
  • The research showed that misjudgments increased as time passed, revealing that continuous accuracy checks are necessary to automatically close alerts using AI alone.
  • The method combines a classification model that reads the alert and outputs a verdict plus reasoning with a second calibration model that reads the alert, judgment and reason and calculates the probability the answer is correct; high-confidence results become candidates for automated processing or priority investigation and low-confidence results are reviewed by human analysts.
  • Training used 388,336 items drawn from 8 consecutive weeks of Windows endpoint alerts.
  • Tuning used 59,162 items from the following 2 weeks.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

CrowdStrike published research on 2026-08-17, titled "Teaching AI to Reason Through Detection Triage," on whether a model can decide if a Windows endpoint alert is a real attack or a harmless false positive [1]. The finding that matters for anyone planning to wire this into a queue is that misjudgments increased as time passed, which means auto-closing alerts with AI alone requires continuous accuracy checks [2].

The design is two models, not one. A classifier reads the alert and produces a verdict plus a reason; a separate calibration model reads the alert, the verdict and the reason, and outputs the probability that the answer is correct. High-confidence outputs become candidates for automated handling or priority investigation, and low-confidence outputs go to analysts [3]. The input is not just a detection name: process, parent and grandparent process, command line, file name and path, how often the file is seen in that customer environment and across all customer environments, sensor severity, action results, MITRE ATT&CK classification, prior analyst write-ups, and a flag for missing fields [13]. Reported components include Nemotron-3-Nano-30B and Nemotron-3-Super-120B, with GEPA, AdaSTaR, LoRA and GRPO [12].

The data split is the interesting part. Training used 388,336 alerts across eight consecutive weeks, tuning used 59,162 from the following two weeks, and the final test used 42,686 from one subsequent week [4][5][6], roughly 490,184 labelled alerts in total [1], with past human true-positive and false-positive calls as ground truth [7]. Overall accuracy was 82.6% [8], so about 17.4% of alerts were classified wrongly before any confidence gating [7]. High-confidence "attack" verdicts hit 98.9% precision at 53.0% recall [9]. High-confidence "no problem" verdicts hit 90.8% precision at 64.8% recall [10].

Then the decay. Precision on the "no problem" class dropped significantly between the tuning period and the final test period: the pre-tuning aim was about 98%, and the later period delivered 90.8% [11], a fall of roughly 7.2 percentage points [2]. Read as an operations number, that is about 9.2% of confidently auto-closed alerts being wrong, close to one in eleven [3]. And this showed up inside the three weeks immediately following the training window [6]. Not a year. Three weeks. CrowdStrike names the mechanism: distribution shift, where live alerts diverge from training data because of new attacks or product updates [14].

The recall figures also cap the savings. At 53.0% recall on high-confidence attacks, 47.0% of genuine attacks do not clear the confidence bar and still need a human [4]. At 64.8% recall on the benign class, 35.2% of false positives keep landing in analyst queues [5]. So the workload reduction is partial by construction, while the risk concentrates in the one action that is hard to reverse: closing something that was real.

The research is explicit that false-positive rate is measured after operations begin and thresholds are adjusted accordingly [15]. That is the actual deliverable. You are not buying a triage robot; you are taking on a recurring measurement obligation, with labelled ground truth, a shift detector, and someone who owns the threshold.

Worth watching: whether vendors shipping auto-close features publish a decay curve alongside their launch precision, and how often they re-tune. Also whether the "no problem" class gets a separate, tighter threshold than the attack class, given that the two error types cost very differently.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories