Skip to content

Product1 publisher3 min readPublished

Testing teams adopted AI for writing tests 3.5 times as often as for judging risk

PractiTest puts AI use across testing organizations at 76.8%, with test creation at 69.6% and risk identification at 19.9%. Sonar has 42% of committed code AI-assisted and heading to 65%, all of it flowing toward the same reviewers.

The Product Desk · Product desk

What happened

  • In the same report, AI use for risk identification, the work of deciding what is worth testing at all, sits at 19.9%.
  • GitLab's Harris Poll of 1,528 developers and technology buyers across six countries found 85% agree the bottleneck has moved from writing code to reviewing and validating it.
  • GitLab also found 87% confident they could identify AI involvement in an incident within 24 hours, while 34% of organizations that had actually had one could not make the call.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • cost Review gets more expensive per item at the same time as it gets more items: 38% say AI code takes more effort to review than a colleague's, and the people paying that are the reviewers who were already the constraint.
  • decision The next quality purchase is really a headcount question, since a review tool cannot decide what it blocks, what false positive rate a team tolerates, or who overrides it the day before a release.
  • constraint Any dashboard built on deployment frequency will keep reporting improvement, which leaves teams arguing for verification capacity without a number that shows the strain.

PractiTest's 2026 State of Testing Report has test case creation at 69.6% and script maintenance at 59.6%, while use of AI for risk identification, the work of deciding what blocks a release, sits at 19.9% [2][3][4].

Divide 69.6 by 19.9 and you get 3.5: adoption ran three and a half times deeper into producing test artifacts than into deciding which risks those artifacts should cover [1]. Teams often describe this as AI taking the grunt work off the tester so the tester can think more. The numbers show grunt work getting faster, and thinking work still landing on the same person, now with more of it.

The split fell this way because test creation and script maintenance produce an artifact you can score in a demo, not because teams are lazy about risk work: a suite that runs, a broken selector that gets fixed. Risk identification produces a decision about what not to test, which has no artifact and no scoreboard until something ships broken. Tools get built where the demo can be graded. PractiTest's cross-section is a count of usage, not of intent, so it does not tell you whether risk work stayed human because teams protected it or because nothing on offer does it convincingly. Either way the queue looks the same.

The inflow keeps climbing. Sonar's survey of more than 1,100 developers put AI-generated or AI-assisted code at 42% of committed code, with respondents projecting 65% by 2027, which is 1.55 times the current share [5][6][2]. In the same survey, 96% said they do not fully trust AI-generated code while 48% said they always verify it before committing, a 48 point gap between stated distrust and stated practice [7][8][3]. And 38% said reviewing AI-generated code takes more effort than reviewing a colleague's [9]. More volume, at a higher unit cost, arriving at the station that was already the constraint. GitLab's Harris Poll of 1,528 developers and technology buyers across six countries found 85% agreeing the bottleneck moved from writing code to reviewing and validating it [10].

The confidence numbers show a similar mismatch. GitLab found 87% confident they could identify whether AI-generated code was involved in an incident within 24 hours, but among organizations that had actually had such an incident, 34% could not make the call [11][12]; 43% said they cannot reliably tell AI-generated code from human-written code at all [13]. Perforce's 2026 State of DevOps Report, drawing on more than 800 IT professionals, has 77% expressing confidence in their AI outputs against 39% maintaining fully automated audit trails [14][15]. As devops.com puts it, the divide is between teams that know which of the two they are and teams that find out during a postmortem [20].

A 2x2 you can fill in this week: one axis is the share of your commits that are AI-assisted, above or below Sonar's 42% [5]; the other is whether a named person owns what the review gate blocks and what false positive rate you will tolerate before people start clicking past it. High share with nobody owning the gate is the quadrant where deployment frequency looks best, because deployment frequency measures the half that got cheaper [19]. The World Quality Report's 17th edition found 15% of organizations have scaled generative AI in quality engineering enterprise-wide against 43% still experimenting, so most readers are somewhere in the middle of this [16][4]. devops.com's own answer is that verification capacity is a staffing question before a tooling one, often a dedicated engineer who owns the pipeline and the gates [18]. The number that would show whether you are right is rework: how much cleanup follows a release [19].

What to watch

  • Whether PractiTest's next State of Testing edition moves risk identification off 19.9% or leaves it flat while creation adoption saturates.
  • Whether Sonar's projected 65% AI share of committed code shows up in 2027, and whether the 48% who always verify moves with it.
  • Whether GitLab's 34% attribution failure among organizations with real incidents narrows once audit trail coverage gets past Perforce's 39%.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories