Security1 publisher3 min readPublished
Gartner: without structured evaluation, only about one in five AI SOC pilots seen achieving measurable gains by 2028
The forecast describes what happens without structured evaluation, and the survey data published alongside it points at the mechanism: teams scope proofs of value against alert queues where 28% of alerts are never opened.
The Watch · Security desk

What happened
- Gartner analysts Craig Lawson and Andrew Davies forecast that by 2028, 70% of large SOCs will pilot AI agents for Tier 1 and Tier 2 operations, but only 15% will achieve measurable improvements without structured evaluation.
- Prophet Security's State of AI in Security Operations 2026 reports 40% of security teams using AI daily, 56% evaluating or piloting it, and 4% with no plans to adopt.
- The same survey found up to 40% of organizations have switched off specific detection rules because they did not have the capacity to work the output.
Compiled by The WatchSomething wrong?How this is made
Why it matters
- exposure An intruder's activity that lands in the unread 28% of alerts stays unreachable to a pilot benchmarked on current work, because the agent is scored on the 72% of alerts the team already opens.
- decision Following Gartner's advice means re-enabling rules that were switched off for capacity before the proof of value starts, and absorbing that alert volume during the test rather than after the purchase.
- constraint The public speed number for AI SOC agents is self-reported by survey respondents, which fails the benchmark test Gartner sets: comparable environment, and production rather than proof of concept.
- precedent With 72% of AI users having already tried building their own tooling, the durability question Gartner poses to vendors applies to in-house builds too, and no one asks it of themselves in a procurement cycle.
Read the 70 and the 15 as shares of the same population and the ratio is what a buyer is being told: 15 divided by 70 is about 21 percent, roughly one pilot in five clearing the bar, with 55 percentage points between running an AI agent and showing anything for it [1][2]. Gartner's Craig Lawson and Andrew Davies attach that gap to the absence of structured evaluation rather than to the agents [1].
The evaluation problem starts before an agent is installed. In Prophet Security's State of AI in Security Operations 2026, respondents said 28 percent of alerts are never investigated, and up to 40 percent of organizations have switched off specific detection rules for lack of capacity to work them [5][6]. A proof of value scoped against current work therefore measures the 72 percent of alerts the team already triages [3]. The queue nobody opened appears in neither the before nor the after. Gartner's first set of questions tells buyers to map their own bottlenecks instead of a vendor feature list, and the write-up adds the corollary: alerts suppressed for capacity reasons may be worth switching back on for the exercise [9].
What that omission has already cost is in the same survey. About 60 percent of respondents said an alert they missed or never investigated turned out to be material, and 34 percent said it happened three or more times in the past year, which puts more than half of the teams with a material miss in the repeat column [7][4]. The survey write-up draws the line between tuning out a rule that has never produced an escalation, which is detection engineering, and switching one off because its output went unread, which is a coverage decision [15].
On outcomes, Gartner's list is mean time to detect, mean time to respond, false positive reduction, and mean time to contain, with containment as the end goal because that is where risk drops [10]. It also tells buyers to ask for benchmarks from comparable environments and to establish whether the numbers came from a proof of concept or from sustained production [11]. Apply that to the one speed figure on offer: 72 percent of survey respondents using AI reported cutting investigation time by at least 25 percent, self-reported, with no stated baseline or measurement window [8]. The durability questions cover general availability, customer base, funding, and how pricing behaves under load, and Gartner accepts that consolidation in the category is likely [12]. That lands on a market where 72 percent of AI users in the survey had already tried building the tooling themselves [13].
Gartner's stated purpose for the question set is separating viable products from the "AI washing" it calls rampant [14]. Its own placement of AI SOC agents moved from Innovation Trigger with single-digit adoption last year to the Peak of Inflated Expectations this year [2][3], which tells a SOC nothing about its own queue. A pilot that reports faster triage without stating how many suppressed rules were re-enabled for the test has measured the existing workload at speed.
What to watch
- Whether Prophet Security's next survey shows the 28% uninvestigated share moving in teams that ran an AI agent pilot.
- Any AI SOC vendor publishing mean-time-to-contain numbers from sustained production rather than a proof of concept.
- Whether triage-focused agents ship detection engineering and threat hunting, the adjacent use cases the write-up expects next.