Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

GitHub is rebuilding its Git storage around the write load agents create

GitHub is rebuilding its Git storage infrastructure as monthly pushes grew 4.9x in a year to 3.35 billion. The strain is on writes, since every push must be durable and consistent before the next agent or CI job can build on it.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying GitHub is rebuilding its Git storage around the write load agents create
Generated illustration

What happened

  • GitHub Actions ran 3.26 billion times in September, more than four times as many runs as a year before.
  • Under the current Spokes system, every repository sits as full copies on the local disks of five fileservers by default.
  • GitHub says thousands of agents on separate branches in one repository create a sustained write rate that converges on a single point in its architecture.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure Teams running agents that commit after every action get GitHub's single-push latency as their throughput ceiling until a new write path ships.
  • constraint Caches and replicas leave write throughput unchanged, so absorbing 4.9x push growth means redesigning how a push becomes durable and visible.
  • decision Teams sending all agent output through one merge queue onto trunk now have reason to weigh which work actually needs to land on that single ref.

GitHub's own figures show writes growing faster than everything else. Total Git activity went from 218.2 billion events a month to 473.3 billion between September 2025 and August 2026 [3]. That is about 2.17x [22]. Pushes grew 4.9x over a year [5], more than double the multiple for activity overall [23]. Commits in September were more than five times the count a year earlier [4]. "Reads are relatively easy to scale: add caches, add replicas, and serve the same bytes to more clients," GitHub wrote [16]. "Writes are way harder." [17]

Today a write lands on Spokes. Spokes keeps a full copy of every repository on the local disks of several fileservers, five by default, and the extra copies give redundancy while spreading reads [8]. When a push updates a reference, a three-phase commit protocol uses a quorum so CI, the web UI and API clients see one consistent repository state [9]. That pairing serves a billion repositories [10]. Five local copies plus a quorum on ref updates is a good answer to durability and read latency at the same time.

Agents change the per-push budget. An agent in a tight loop commits or checkpoints after nearly every action, so its speed is bounded by how fast a single push completes [11]. Latency a person would never notice becomes the limit [11]. Agents also do not save up a morning's work for one push before lunch [11]. Put thousands of agents to work in one repository, each on its own branch, and their writes add up to a constant load aimed at one point in GitHub's architecture [12]. Merges then funnel onto one ref through trunk-based development, release trains and merge queues [13]. Pull request merges are up nearly 4x in a year [6].

The busiest repository took roughly a billion requests in August [2]. Spread evenly, that is about 373 requests a second, every second of the month [24]. It sits at the far end of a distribution where activity climbs sharply, and GitHub says the gap to a typical repository is wider than most people expect [20]. For that figure to describe your repository, you would need what the post finds at the top of the curve: a large team with busy CI pipelines running alongside a growing agent fleet [21].

Reads still multiply. CI and code scanning clone or fetch the same branch tip thousands of times a minute [14]. Actions ran 3.26 billion times in September [7]. Behind both, GitHub compacts repository data and cleans up unneeded objects, and every new write adds to that work [15].

The rebuild, as described, is of GitHub's Git infrastructure: the storage and replication under the repositories [19]. Four of the five pressures the post lists are capacity problems inside GitHub. The fifth, contention on one ref, comes from workflows teams run on top [13].

The available text stops before the replacement design, mid-sentence: "However, the mechanism we use for durability is the same one we use for" [1]. The Spokes paragraph just before it describes one set of five copies that provides both redundancy and read spreading [8]. On the strength of that paragraph, I'd expect the new design to size durability and read capacity separately.

What to watch

  • The rest of GitHub's post or a follow-up describing what replaces five-copy replication and the three-phase commit on ref updates.
  • Any push-latency target GitHub publishes, since an agent's loop speed is bounded by how fast one push completes.
  • Changes to merge queue or ref-update behaviour for the highest-volume repositories, where every merge contends on one ref.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories