Skip to content

Product1 publisher3 min readPublished

Flux reads the commit history to answer whether AI coding spend delivered

The company's five new measures are generally available and pull from merged changes, review timing, security findings and cost classification. Every score is set against the organization's own past performance.

The Product Desk · Product desk

Photograph accompanying Flux reads the commit history to answer whether AI coding spend delivered
Photo: letsdatascience.com

What happened

  • Flux Cyber expanded its platform with tools meant to show engineering leaders whether their investment in AI for software development is paying off, and the company has broken that into five blind spots.
  • One capability, defensible spend, sorts engineering work into capitalizable and operational categories for spending decisions and R&D tax credits, and excludes token usage from the calculation.
  • Calibrate Ventures led a $5 million round in Flux in June, with True Ventures and Glasswing Ventures also investing.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • constraint Because deployment frequency and lead time are measured against the organization's own history, a team with a slow baseline is graded against that slow baseline, and the output supports no claim about how it compares with other companies.
  • exposure Attributing significant changes from commits and pull request history makes visible the engineers whose work never appeared in the ticket tool, and it makes them visible to whoever reads the report, not only to their tech lead.
  • capability Splitting engineering work into capitalizable and operational categories puts a code-analysis tool upstream of a tax position. An auditor may eventually be reading its classifications.
  • contradiction The measurement gap here rests on Flux's own account of its customer conversations, and the DORA finding it leans on supports the claim that adoption proves little; it does not support the claim that code-derived measures prove more.

What engineering leaders send upward, according to Flux, is tickets, story points, sprint velocity and token counts, numbers that describe how busy a team has been. [3] Ted Julian, the company's founder and chief executive, called AI "the biggest bet most engineering organizations have ever made." [6] "Tickets and adoption rates describe intent," Julian said. "The code itself shows what the team actually built and delivered." [7]

Verified velocity sorts every merged change by type and compares deployment frequency and lead time with the organization's own history. [10] Trusted review times how long a change waits for its first reviewer and weighs review depth against the change's size and risk. [11] Auditable work draws on commits and pull request history. [12] Continuous quality follows new and resolved security findings, newly added dependencies and failure and recovery trends against each team's norms. [13] Defensible spend sorts the work into capitalizable and operational categories. [14]

All five inputs read the work; a user is not one of them. [16] A commit is harder to inflate than a story point. The swap on offer is from process measures kept in a ticket tool to process measures pulled from the repository.

Google Cloud's DORA program found in its 2025 report that 90% of software professionals use AI at work. [4][17] So an adoption rate separates a team from about one in ten. The authors wrote that AI "doesn't fix a team; it amplifies what's already there." [5] That applies to measurement as well. A thin-reviewing team that measures review depth learns its review is thin, and the measure adds no reviewers.

The most concrete of the five is the one built for finance. Defensible spend is meant to give leadership something solid to point to for spending decisions and research and development tax credits, and Flux said the measure covers engineering effort and where it goes, with token usage left out. [14]

For the person who has to deploy this: the company says all five run on the code analysis the platform already performs, so there is no instrumentation to add and no change to how teams work. [8] The release leaves out pricing. [18] Calibrate Ventures led a $5 million round in June, with True Ventures and Glasswing Ventures also investing. [15]

Before a pilot, I would ask of each capability which decision a bad number would reverse. Review load concentrated on two senior engineers means a reassignment or a hire. Lead time doubling in the quarter after a coding-assistant rollout means changing the rollout. Flux says each blind spot is assessed separately, because an organization can be solid in one area and exposed in another. [9]

What to watch

  • Named customers with before-and-after lead time or review latency numbers. Those numbers would show whether the five measures change decisions or only reports.
  • Whether Flux publishes pricing and packaging, since the release does not include a price.
  • Whether DORA's next report tests code-derived delivery measures against AI adoption rates.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories