Skip to content

Build1 publisher3 min readPublished

Time to first review quadrupled in the activity spike after Flux's customer adopted AI assistants

Flux gave Lets Data Science a three-month before-and-after from one anonymized customer: deployment frequency improved, static-analysis findings rose, and the company was working without developer-level AI usage data for either.

The Engineer · Build desk

Photograph accompanying Time to first review quadrupled in the activity spike after Flux's customer adopted AI assistants
Photo: letsdatascience.com

What happened

  • Flux CTO Aaron Beals compared one anonymized customer's repositories and deployment history for the three months after it adopted AI coding assistants against the three months before.
  • Deployment frequency and pull-request lead time both improved over that period, according to Beals.
  • During an activity spike in the same window, time to first review rose from 6.4 hours to 25.6 hours, a fourfold increase in the wait before review begins.
  • Static-analysis findings for security, complexity and technical debt rose, along with the number of changes classified as bug fixes and refactoring.
  • CEO Ted Julian's illustrative cost example assumes 20 engineers, $30 per person a month for subscriptions and $40 for API usage, putting direct tool costs at $1,400 a month.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Flux's comparison was period against period, and it ran without developer-level AI usage data for this customer, so the case can point a team at what to investigate while falling short of showing that the assistant lengthened the queue.
  • cost The subscription line is the cheapest part of the bill to count: Julian puts setup, continuing integration work, review, rework and incident handling outside it, and those are paid in engineer hours.
  • contradiction Flux argues review and rework need more attention, while DORA's guidance already specifies deployment rework rate and change fail rate, so a team seeing only delivery speed is under-using a framework it likely already cites.

The added wait before anyone picks up a change is 19.2 hours [23]. If a working day is eight hours, a change that used to sit for most of one day now sits for a bit over three [29]. Four times the wait does not mean four times the work: Beals says the figure is not a measure of the reviewers' labor cost [7].

The number also comes from an activity spike inside the window, not from the three-month average [6]. For it to mean anything on another team, the baseline has to be built the same way, and LDS's recommendation is to write down the start point, the end point, the repositories and the exclusions first, then keep them consistent [16]. For the baseline, Beals points teams at their own previous 12 weeks and away from industry averages [17]. The test he and Julian describe applies to assistants used for software, data pipelines or ML services: follow what happens after code is generated, and establish whether any freed capacity reaches useful work [28].

Where the clock stops changes the answer. DORA defines change lead time from a code commit to production deployment, so a team whose measure ends at PR merge is timing a shorter interval [15].

Before reading anything into the comparison, Beals lists what else could have moved: staffing, work mix, holidays, automated dependency updates, release cadence, and whether scanner rules or repository coverage changed [18]. A static-analysis finding is a signal to investigate; a confirmed production incident is something else [9]. On the rise in fixes he leaves the direction of causality open, because the burst of changes may have introduced the problems, or finding the problems may have prompted the burst of repair [11]. "It's strong enough to act on. It isn't proof," Beals said [13].

Seventy dollars per engineer a month for subscriptions and API usage works out to $16,800 a year across the 20-engineer example [25][24]. Julian keeps setup, continuing integration work, review, rework and incident handling separate from that invoice [20]. The example leaves those hours unpriced. "Review deserves more attention than it gets," Julian told LDS [21].

DORA's current guidance already includes deployment rework rate and change fail rate, and advises interpreting metrics in the context of an application or service [14]. Against that, the interview is no evidence that established delivery measures ignore rework [26]. A chart showing more fixes after adoption cannot on its own establish that the assistant created the defects, since a migration, a change in release practices or different work assignments would show up the same way [27]. What the customer case adds is one observational comparison, useful for deciding what to investigate [19]. Settling what the tool caused is beyond it. Flux's account is company-reported, and LDS did not verify it independently [3].

What to watch

  • Whether Flux publishes a case with per-developer AI usage tied to individual changes rather than periods around adoption.
  • Whether the customer's time to first review fell back toward 6.4 hours once the activity spike passed.
  • Whether the static-analysis findings stayed elevated after the single week that carried most of the increase.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories