Skip to content

Build1 publisher3 min readPublished

Uber's SubmitQueue tests pending changes against each other to keep six monorepos green

Uber's SubmitQueue builds a speculation tree of pending changes to keep mainline green in monorepos that take more than 65,000 changes a month. At that volume, a one-at-a-time queue would need every build done in about 40 seconds.

The Engineer · Build desk

Illustration accompanying Uber's SubmitQueue tests pending changes against each other to keep six monorepos green

What happened

  • In Juloori's failure case, two changes branched from the same commit each pass CI, the first merges cleanly, and the second breaks the mainline build.
  • Uber has about 4,500 engineers in more than 10 development centres committing mainly to six monorepos.
  • Those monorepos hold about 135 million lines of code and feed more than 100,000 deployments a month.
  • Juloori listed delayed rollouts, developers unsure why their CI runs fail, and harder rollbacks as the costs of a broken mainline.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint A green branch build certifies a change only against the commit it branched from, so a busy monorepo's merge gate has to test the combination that will actually land.
  • decision Choosing a merge queue starts with measuring change rate against build time, because speculation only pays for itself once a serial queue falls behind.
  • precedent If coding agents raise commit concurrency as Juloori anticipates, repositories far smaller than Uber's will reach the point where a serial queue stops keeping up.

The simplest merge queue takes one change, rebases it on the current tip, builds it, lands it if it passes, and moves to the next. That alone fixes the case Dhruva Juloori of Uber's developer platform opened his talk with, where two developers, he said, had "tested their changes individually but not collectively together" [1][15]. It also limits the queue to one build at a time.

Uber's monorepos take more than 65,000 changes a month [5]. Spread evenly over 30 days, the rate comes to about 2,170 changes a day [1]. A serial queue would need each build to finish in roughly 40 seconds to keep pace [2]. That figure assumes changes arrive at an even rate around the clock. Any slower build and the queue grows all day.

SubmitQueue builds ahead. Its enumerator takes the queue of pending changes and builds a binary decision tree of all the possible builds, which Uber calls a speculation tree [10]. A tree that splits in two for every pending change doubles with each one added, so ten pending changes give 1,024 possible paths [3]. The next two components decide which of those builds are worth running. The profiler ranks changes in landing order from predicted build times [11]. The prioritizer computes, for each node, the probability that its build will be needed, by predicting whether each change will succeed [12]. The fourth component, the selector, is in Juloori's words "the most easy one" [13]. The published transcript cuts off one sentence later.

Speculation spends CI capacity on predictions. Builds on a correctly predicted path let a change land as soon as they finish. Builds on a mispredicted path are thrown away. Juloori said SubmitQueue "utilizes CI resources very efficiently" and that "it always guarantees green mainlines at scale even with thousands of changes per day or hundreds of changes per hour" [8][9]. Both statements are Uber describing its own system. The transcript does not include a prediction hit rate, a wasted-build share, or the SLO values behind the promise of "reasonable SLOs" [8].

Uber's figures describe Uber's workload [3]. The break Juloori described needs only two concurrent changes, so it can hit a team of any size [2]. Scale decides which fix is affordable. In my view, a team needs speculation when its daily change count multiplied by its typical build time exceeds the hours in which changes arrive. Below that point a serial queue keeps main green without any prediction step. I think the speculation tree is good engineering for Uber's volume, and more machinery than a repository with a few dozen changes a day needs. Juloori raised agents as the next source of concurrency: "when we have an army of agents concurrently committing to a single codebase, it gets very hard to maintain the stability of the mainline," he said [14].

What to watch

  • Uber publishing SubmitQueue's prediction hit rate or wasted-build share, the figure that would test its CI-efficiency claim.
  • The actual SLO values behind SubmitQueue's promise of reasonable land times, and how they degrade when the prioritizer's predictions miss.
  • Commit-rate data from teams adopting coding agents, showing whether they cross the change-rate-times-build-time limit of a serial queue.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories