Skip to content

Product1 publisher3 min readPublished

Faros telemetry shows pull requests nearly quadrupling across 22,000 developers in two years

Two years of compounding growth in pull request size lands on the same reviewers, while Futurum's adoption figures put AI at 40.2% in code generation against 6.2% in deployment decisions. That gap is where the queue forms.

The Product Desk · Product desk

Photograph accompanying Faros telemetry shows pull requests nearly quadrupling across 22,000 developers in two years
Photo: faros.ai

What happened

  • Faros AI telemetry covering roughly 22,000 developers found pull requests grew 154% larger in 2025 and a further 51.3% larger in 2026.
  • The same telemetry recorded developers handling 67.4% more pull-request contexts per day, so larger changes are arriving across more simultaneous threads of work.
  • A LangChain survey of more than 1,300 practitioners found observability implemented by nearly 89% of respondents but evaluation by only 52%.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • constraint Buying more generation capacity does not add review hours, so a 3.84x larger package consumes a review budget that no seat purchase increases.
  • decision Teams renewing AI coding tools now have to choose between more generation seats and the pre-merge gates that let the output through, because the 6.2% figure says almost nobody has bought the second.
  • contradiction devops.com argues review capacity failed to keep pace, but the telemetry it cites measures batch size and context load rather than wait time, so the bottleneck is inferred rather than counted.

The person who pays for the 154% is whoever opens the diff on Monday. Compound the two years and you get 2.54 times 1.513, or 3.84 [1]: a change that used to run 400 lines would now arrive at about 1,540 [2]. The same reviewer is also carrying 67.4% more pull-request contexts in a day [2], which is the part that does the damage, because switching is what turns a large diff into something skimmed rather than read.

The story a rollout tells itself is that ten times the generation capacity becomes ten small, reviewable changes. The devops.com read of the same telemetry is that without a change in delivery discipline it becomes one much larger pull request instead [14], and the adoption numbers say discipline is where nobody has been spending. Futurum figures, cited in Techstrong's report The Great Unification, put AI use at 40.2% in code generation and 37.7% in review, then 13.2% in CI/CD operations and 6.2% in deployment decisions [3][4][5][6]. Generation adoption runs about 6.5 times deployment adoption [3]. And 9.5% of organisations release daily, even though 58.3% place their own DevOps maturity at standardising or mastering [7][8].

That 6.2% carries two defensible readings, and devops.com offers both: humans deliberately kept at the highest-consequence boundary, or delivery machinery that cannot absorb what upstream now produces [12]. Either reading puts the constraint in the same place.

The LangChain survey of more than 1,300 practitioners makes the second one harder to wave off. Observability sits at nearly 89% and evaluation at 52%, a 37-point gap [9][10][4], and many of those evaluations run offline rather than as a gate that can stop a release [11]. So a team can describe an agent misbehaving in production but cannot show before merge that a prompt or model change did not make things worse. With AI-dependent code the same input can produce different outputs, and a model can change behaviour with no application-code change at all [16], so the pass/fail contract CI was built on stops holding. Rollback also gets slower and dearer, since model artifacts are large and the failure signal is a statistical drop in answer quality rather than a spike in HTTP 500s [15].

The telemetry does not show the queue itself. Faros reports package size and context load; the claim that review capacity failed to keep pace is an assertion in the devops.com piece rather than a published review latency, merge time or defect-escape figure [13][5]. Anyone deciding on next year's seat count is reasoning from batch size and adoption gaps, not from a measured wait.

The forcing function is cheap enough to run without buying anything. Two axes: whether AI writes this class of change, and whether a gate can stop that class of change before merge. Three of the four boxes are survivable. The fourth, heavy generation with no pre-merge gate, is where a 3.84x diff turns into an escaped defect, and 6.2% suggests it is the crowded box [6][1]. The two numbers that tell a team it is sitting there are the median time from pull request open to first substantive review comment, and the share of merges that arrive with an evaluation result attached. Both are already in the systems most teams run.

What to watch

  • Whether Faros publishes review latency or merge time alongside pull request size, which would show whether a queue is actually forming.
  • Whether the next Futurum cut moves the 6.2% deployment-decision figure, and whether it moves because gates got automated or because humans stepped back.
  • Whether evaluation adoption above 52% in the next LangChain survey shows up as pre-merge gates rather than offline dashboards.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories