Skip to content

Build1 publisher3 min readPublished

Databricks moves feature serving to streaming, and the 200ms is measured to the write

A p99 of 200 milliseconds from Kafka to the online store beats batch cadence by a wide margin. It also stops before your model reads the feature.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Databricks Feature Store lets you author a feature once and use it everywhere: the same definition drives large-scale batch flows offline and highly fresh feature pipelines online.
  • Spark pipelines are an established way to process bulk data in the Lakehouse for historic baseline features; running these batch jobs on a regular schedule is well understood but introduces minutes to hours of lag. Databricks says that for baseline signals this lag is an acceptable price for simpler infrastructure.
  • Databricks states that when models require fresh signals the batch infrastructure breaks down, and that getting down to seconds or milliseconds is not possible in existing feature store platforms.
  • Databricks says that to deliver fresh features today, data scientists are forced to implement complex, streaming-specific logic to handle aggregations and to stand up custom hosted infrastructure.
  • The framework orchestrates Spark Real-Time Mode (RTM) for continuous stream processing, Lakebase for streaming-optimized online storage, and Model Serving for retrieval at scale.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Databricks has published the internals of a streaming path for its Feature Store, in which a definition authored once drives both large-scale offline batch flows and a fresh online pipeline [1]. For teams who built a bespoke low-latency serving tier because scheduled Spark was too slow, the useful detail is not the freshness claim but where the measured latency boundary sits [7].

The problem statement is familiar. Scheduled Spark jobs over Lakehouse data are well understood for historical baseline features, but batch cadence introduces minutes to hours of lag [3]. Databricks argues that lag is an acceptable price for baseline signals and that the infrastructure breaks down when models need fresh ones, asserting that getting down to seconds or milliseconds is not possible in existing feature store platforms [3][4]. That is a vendor claim about competitors and should be read as one. Its description of the workaround is harder to argue with: data scientists writing complex streaming-specific aggregation logic and standing up custom hosted infrastructure [5].

The replacement, according to Databricks, orchestrates Spark Real-Time Mode for continuous stream processing, Lakebase as streaming-optimized online storage, and Model Serving for retrieval at scale [6]. The published number is an end-to-end p99 of 200ms, from an event arriving in Kafka to availability in the online feature store [7]. Note both ends of that measurement. Kafka is still upstream in Databricks' own account of the path, so the queue you run does not go away [7]; and the 200ms covers ingest through write-visibility, not the model's read of the feature or the inference that follows. Against a one-minute batch cadence, the fastest end of the stated minutes-to-hours range, 200ms is about 300 times tighter [8]. Against an hourly job it is about 18,000 times [9].

Mechanically it is the design most teams converge on independently. Each transaction event, carrying amount, location, user id and merchant details, is routed to a stateful pipeline that consults a local RocksDB instance holding the user's running total, with expiry times bounding the window to the last 10 minutes, increments the value locally, then writes it to Lakebase [10]. The 10-minute sum is then fetched alongside the user's 30-day purchasing baseline, and a sum well above that baseline is what the model reads as potential fraud [11][12].

The cost question lives in the window semantics, and Databricks is reasonably candid about it. Feature Store supports three time window types [13]. Tumbling and sliding windows emit fewer updates, are cheaper to maintain, and fit naturally into simpler scheduled pipelines [14]. Rolling windows trade that efficiency for maximum freshness, where every new event immediately affects the value served [15]. One online write per event per key is not a rounding error at fraud-traffic volumes, and the post does not price it.

Worth watching: whether the 200ms holds once Model Serving retrieval and inference are inside the measurement, given that Model Serving is a separate component in the chain [6][7]; and what the state layer costs when rolling windows are the default choice for anything latency-sensitive [15]. Anyone running a Redis tier purely to close the batch-lag gap now has a narrower justification than they did, but the number to compare is total read-path budget at their own key cardinality, not the write-side p99 in the blog post [7].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories