Skip to content

Build1 publisher3 min readPublished

Uber's ServiceScale lets several orchestrators scale the same Kubernetes workloads

Uber built ServiceScale so its failover orchestrator and its deployment controller can both scale workloads across a fleet of 3 million cores. During an outage, low-tier services shrink to make room for high-tier ones, a plan meant to replace idle reserve capacity.

The Engineer · Build desk

Illustration accompanying Uber's ServiceScale lets several orchestrators scale the same Kubernetes workloads

What happened

  • Uber's engineers rejected adding failover logic to UDC, their deployment controller, because it already sits on the hot path for service lifecycle operations.
  • Each orchestrator records its scaling desire in a ServiceScale custom resource, and a new Service Scale Controller reconciles the combined intent into Kubernetes objects.
  • Steady-state and temporary failover scale both live in the ServiceScale spec, so failback does not require reconstructing state from logs.
  • Kubernetes v1.36, released in April 2026, added a comparable check under which a controller does not act if its cache lags the resource version it wrote.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Idle reserve stops being a platform line item and becomes a cost that low-tier service owners pay, since their workloads shrink during another region's outage.
  • constraint Any controller whose status can trigger an irreversible step has to prove its cache has seen its own write before reporting success, or a few seconds of lag can advance a workflow early.
  • decision Platform teams adding a second scaling writer now have a worked alternative to branching their deployment controller: a separate intent object with its own reconciler.
  • precedent Once controller-runtime ships read-your-own-write semantics, guardrails like Uber's generation annotation become default behaviour for controllers built on it, and homegrown versions become code to delete.

Uber's fleet launches 1.5 million pods a day [3], about 17 a second [1]. UDC reconciles the intent service owners set in Up into Kubernetes primitives for that fleet [4]. Putting failover code inside it would have put a path exercised only during outages next to code that every lifecycle operation depends on [7]. I think the team made the right call. Grishechko and Paruchuru wrote that "a regression in failover handling wouldn't stay isolated to failover" and could affect normal deployments across the fleet [8].

Keeping intent inside Kubernetes is the choice I would copy. The authors listed what they declined to build: "We didn't want an additional external database, a separate coordination service, or a control plane that'd become harder to debug under incident pressure," they wrote [10]. An external coordination store is one more system that can be down during the outage it exists to handle. With intent in a custom resource, an engineer who sees a wrong replica count reads the ServiceScale object and learns which orchestrator wanted what [11].

InfoQ describes the controller as letting several orchestrators manage the same workloads safely [1]. The production lessons in the post show where that safety came from [19]. The first was stale reads. Uber's controllers consume resources through informer caches that can lag reality by a few seconds [13]. Up treated a status field as terminal input, so a success signal from UDC could trigger the next irreversible step in a workflow [13].

The guardrail is small. When a controller updates downstream resources, it attaches its current generation as an annotation. It then verifies that its cache reflects at least that generation before it reports status [14]. The Kubernetes project is working with controller-runtime to give every controller built on that tooling the same read-your-own-write semantics [16]. Software engineer Prasad M K, writing on LinkedIn, called the gap "an API contract problem, not a backend cache problem" and recommended version tokens to validate reads against writes [17].

The second lesson came from multiple writers. It surfaced when UDC and SSC began updating at the same time [18]. Splitting intent from execution creates that condition by design, so any team adopting the pattern has the problem from its first failover.

The capacity claim transfers only under specific conditions. Uber's plan frees room by scaling low-tier workloads down and high-tier workloads up during a failover [6]. That replaces the reserved idle capacity it historically kept in every data centre [5]. For the swap to cover a regional outage, the surviving region must hold enough low-tier compute to absorb the rerouted high-tier load. Those low-tier services must also tolerate shrinking while another region is down. A company whose regions run mostly critical services has little to reclaim. The reported account sizes the fleet [3] but does not say how much reserve capacity Uber has retired.

What to watch

  • A figure from Uber for reserve capacity retired, or results from a regional failover drill run through ServiceScale.
  • The controller-runtime read-your-own-write change landing, and whether Uber drops its generation-annotation guardrail in favour of it.
  • How SSC resolves conflicting intents once more than two orchestrators write to the same ServiceScale object.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories