Skip to content

Build1 publisher3 min readPublished

CrowdStrike's 98% and Simbian's 3.8% are measuring different SOC jobs

Microsoft and Omdia put the average security operations centre at close to 3,000 alerts a day with 42% never investigated, while a 2026 benchmark scored the best frontier model at 3.8% on finding malicious events in raw Windows logs.

The Engineer · Build desk

Illustration accompanying CrowdStrike's 98% and Simbian's 3.8% are measuring different SOC jobs

What happened

  • Microsoft and Omdia's State of the SOC 2026 puts the average SOC at close to 3,000 alerts a day, with 42% never investigated and 46% turning out to be false positives.
  • A 2026 Simbian Research benchmark gave frontier models raw Windows logs with no hints, and the best model correctly flagged malicious events 3.8% of the time.
  • Gartner projects 70% of large SOCs will pilot AI agents for Tier 1 and Tier 2 work by 2028, with only 15% seeing measurable improvement absent a structured evaluation process.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • contradiction Palo Alto declared 2025 the year of the autonomous SOC while Gartner's own research says there will never be one, so a buyer is funding a category whose suppliers and analysts do not agree on the end state.
  • exposure Auto-closing alerts moves the false-negative liability from the analyst who never opened one to whoever set the agent's confidence threshold, and that person is usually named in the runbook, not the contract.
  • decision Gartner ties measurable improvement to a structured evaluation process, which puts the harness in the pilot's scope of work and gives the buyer something to withhold payment against.
  • cost If the 40-plus hours a week CrowdStrike reports removing holds, the saving lands in headcount planning as roughly one analyst post, so it has to be argued with the people who own hiring.

Three thousand alerts a day, spread evenly, arrive 28.8 seconds apart [2]. CrowdStrike's 2026 Global Threat Report puts average eCrime breakout time at 29 minutes, and the fastest it observed at 27 seconds [6]. So about sixty alerts land inside one average breakout window [3]. Of the day's queue, 42% never gets opened, which at that volume is roughly 1,260 alerts [2][1].

The two numbers being published about triage agents measure different work. CrowdStrike's Charlotte AI Detection Triage operates on detections, and the company reports over 98% agreement with human expert decisions under what it calls bounded autonomy, plus 40-plus hours of manual work removed a week [7]. A detection is a candidate that some rule already decided was worth raising. The Simbian Research benchmark handed frontier models raw Windows logs with no hints: the best model correctly flagged malicious events 3.8% of the time, and no model cleared a 50% recall bar across MITRE ATT&CK tactics [16]. For the 98% to transfer to your shop, the candidate set has to come from detection content you already trust. If the job is finding what no rule fired on, 3.8% is the closer measure.

Simbian's Alankrit Chona, Igor Kozlov and Ambuj Kumar wrote that current LLMs are "poorly suited for open-ended, evidence-driven threat hunting despite strong performance on curated Q&A security benchmarks" [17]. Splunk's 2025 State of Security put the share of leaders who fully trust AI for mission-critical tasks at 11% [18].

Gartner titled its research "Predict 2025: There Will Never Be an Autonomous SOC" and argues that people will always contribute key capabilities, so leaders should aim AI at augmentation [13]. It projects that by 2028, 70% of large SOCs will pilot AI agents for Tier 1 and Tier 2 work, and that only 15% will see measurable improvement without a structured evaluation process [14]. Read both percentages as shares of large SOCs and about one pilot in five pays off [4]. AI SOC agents also reached the Peak of Inflated Expectations on Gartner's Hype Cycle in a single year [15]. One year is a quick trip.

The input side is where the design gets uncomfortable. Prompt injection sits at number one on the OWASP Top 10 for LLM applications because these systems cannot reliably separate trusted instructions from untrusted data, and a triage agent's entire job is ingesting untrusted data [19]. MITRE ATT&CK already catalogs alert flooding as a defense-evasion technique [20]. Flooding a queue humans cannot clear buys an attacker analyst-hours. Against a queue an agent closes automatically, I would expect the same technique aimed at the confidence threshold.

IBM's 2025 Cost of a Data Breach Report recorded the first decline in five years, a 9% drop in average breach cost to $4.44M and a breach lifecycle down to a nine-year low of 241 days, which it attributed to faster AI-driven detection and containment [11]. The same report puts the saving for organizations using AI and automation extensively at close to $1.9M per breach against those using none [12]. That is a comparison between two populations of organizations, so it carries whatever else distinguishes a heavily automated security team from one with no automation at all.

The labour case does not depend on any of that. Tines' survey of 468 analysts found 71% experiencing burnout, 64% considering leaving within the year and 69% saying they are understaffed [4], against an average analyst tenure of roughly 18 to 24 months [5]. Closing the repetitive work on the 46% of alerts that turn out to be false positives is a staffing decision with a measured denominator [3]. The open-ended hunt is the part the benchmark scored at 3.8% [16].

What to watch

  • Whether Simbian reruns its benchmark with vendor detections as input instead of raw logs, which is the setup vendor numbers are measured in.
  • Whether Gartner's 15% measurable-improvement figure moves once evaluation harnesses ship with the agents instead of being built by the buyer.
  • Whether any vendor publishes prompt-injection resistance numbers for a triage agent reading attacker-controlled log fields.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories