Skip to content

Build1 publisher2 min readPublished

Cloudflare K2 runs a partitioned event log on R2 with about a second of p99 produce latency

Cloudflare opened K2, a serverless event log built on R2 object storage, to public beta at about one second of p99 produce latency. For that wait, teams get producers decoupled from consumers and no Kafka cluster of their own to run.

The Engineer · Build desk

Illustration accompanying Cloudflare K2 runs a partitioned event log on R2 with about a second of p99 produce latency

What happened

  • Producers write to a stream stored as an ordered log, and consumers can either split the reads among themselves or each receive every message.
  • Cloudflare first built K2 as the durable ingest buffer for Basin Pipelines, whose pull-based engine needs events stored somewhere before it reads them.
  • Cloudflare says its edge, spread across more than 335 cities, often cannot run conventional distributed systems software such as Kafka.
  • K2 accumulates incoming writes in memory on an edge service, then writes each batch to R2 as a single segment file.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure R2 availability becomes a dependency for every K2 producer and consumer, because K2 delegates durability, replication and offset ordering to it.
  • capability A team can add another independent reader of the same events without resizing producers, and a reader that goes down for a long stretch catches up from retention.
  • decision Async delivery on Cloudflare is now a three-way choice among Queues, Basin Pipelines and K2, and a team on Queues has to decide whether an ordered, retained log is worth a second of produce tail.

Cloudflare's stateful services get relatively small slices of machines, those machines are relatively ephemeral, and the network between them is often the public Internet [8]. K2 keeps its partitioned log in R2 [4] and hands replication and consensus to the storage layer [10]. Cloudflare says R2 offers eleven nines of durability and strongly consistent APIs [9]. I think handing consensus to storage is the right call for this fleet. Keeping a quorum healthy among short-lived nodes on public links is the hardest part of operating a distributed log, and K2 has no quorum of its own to keep [10].

Ordering comes from R2 as well. K2 gets strictly incrementing offsets from R2's atomic operations, with no separate coordination service [13]. Compute and storage scale independently, and Cloudflare says that keeps large volumes of history cheap to store [10].

Customers are also getting a buffer Cloudflare already uses to keep a promise of its own. Pipelines commits to never dropping an event once it has been accepted [6].

The cost is on the write path. R2, like other object stores, cannot append [11], an awkward property for the backing store of a log. Each segment has to be large enough to cover the cost of writing and reading it [11], and Cloudflare describes the wait for a batch to fill only as "a short period" [12]. In the initial release, that wait plus an object-store write slower than local disk comes to about one second of produce latency at the 99th percentile [14]. Both waits sit inside the produce call, so I take an accepted event to be one already written to R2 [1]. Cloudflare does not state the acknowledgement point in those words.

That second is Cloudflare's measurement of its own first release [14]. It covers produce only. Time to a consumer adds the read path on top, and Cloudflare says more design detail is coming in a technical deep dive [15].

Cloudflare's own example is an ecommerce backend whose completed transactions feed an analytics system and a fraud detection service [17]. With direct RPC, an overloaded or unavailable consumer means dropped events, according to Cloudflare [18]. Analytics can absorb a second of produce tail. A fraud check that holds a checkout until its event is accepted would pass that second to the customer [14].

What to watch

  • Cloudflare's promised technical deep dive on K2, particularly the length of the batch window and the exact point at which a produce is acknowledged.
  • Whether the roughly one-second p99 produce latency falls in releases after the initial beta.
  • Pricing and retention limits for K2 streams when the service leaves public beta.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories