Build1 publisher3 min readPublished
Jumio's sub-100ms feature store is mostly a consistency build, not a speed build
AWS published the architecture behind Jumio's real-time feature store. Three of the four problems it starts from are about duplicated definitions and hand-copied code, not milliseconds.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Jumio is an identity verification provider that helps businesses detect fraud and build digital trust, and its ML models needed a real-time feature store to provide these services in real time.
- The AWS post uses Jumio's case study to show how to build a real-time feature store, states the pattern applies to ML use cases requiring sub-100ms latency for real-time predictions, and says it shows the architecture, the design trade-offs, and their impact on Jumio's workload.
- Of the four pre-existing issues AWS lists, three (data duplication, manual production deployment, delayed event handling) are consistency or process problems and one (latency challenges) is a speed problem.
- Before building the real-time feature store, feature engineering and deployment at Jumio were often fragmented and inefficient.
- Data duplication: teams maintained their own offline feature stores, resulting in redundant data and inconsistent feature definitions.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
AWS has published the architecture behind the real-time feature store used by Jumio, an identity verification provider that helps businesses detect fraud and build digital trust, presenting it as a pattern for ML use cases that need features served in under 100 milliseconds [1] [2]. What makes it worth reading is the problem list that comes before the diagram: of the four failures named, only one is about speed [3].
According to the AWS post, feature engineering and deployment at Jumio were fragmented and inefficient before the rebuild [4]. Teams each maintained their own offline feature stores, which produced redundant data and inconsistent feature definitions [5]. Features trained offline were then manually re-implemented in production code in Java or Python, which AWS says increased the risk of mismatches and bugs [6]. That is training and serving skew manufactured by a deployment process, and no amount of in-memory storage fixes it. The third non-latency problem is time itself: certain event types arrive with delays or in irregular patterns, appearing shortly after initial activity or several weeks later because of extended review processes [7]. The fourth is the familiar one, fraud detection needing immediate access to features, including outputs from upstream models [8].
The stated requirements track those problems rather than the marketing surface. The platform has to scale to high request volumes and a growing feature catalog while allowing schemas to evolve without disrupting existing workflows [9]. It has to support conditional feature creation and selection based on event times [10]. It has to serve fraud detection features in under 100 milliseconds [11]. The offline store has to ingest in near real time so that retraining, debugging, evaluation and monitoring can work from backfilled history [12]. And it has to let cross-functional teams introduce features independently, with minimal coordination overhead [13].
The build is streaming-first and deployed in three AWS Regions: us-east-1, eu-central-1 and ap-southeast-1 [14]. Events land in Amazon Kinesis Data Streams, Apache Flink applications on Amazon Managed Service for Apache Flink process and enrich them in flight, and the resulting features are written directly to Amazon SageMaker Feature Store, which holds them in memory for inference reads [15] [16]. A parallel path sends events through Amazon Data Firehose to Amazon S3, where an S3 event notification triggers Amazon EMR for the heavier transformations; EMR writes results into SageMaker Feature Store as cold data and into an Apache Iceberg table that serves as the offline store [17] [18].
Two consequences follow from that shape. Because Flink does the computation at ingestion, the sub-100ms path at inference is a lookup, not a calculation, which is why the latency requirement is the tractable half of the problem [19]. And the online store has two writers, the Flink hot path and the EMR cold path, so online and offline agreement rests on both paths producing the same values rather than on one shared implementation [20]. The hand-copied Java or Python has been removed, but the risk it represented moves into keeping two pipelines semantically identical.
Watch for whether schema evolution actually happens without disrupting live workflows [9], since three regions means three copies of every definition [14]. Watch how the EMR path reconciles events that arrive weeks late against features already served [7] [17]. And note this is a vendor-authored case study that promises design trade-offs and their impact on Jumio's workload [2]; the trade-offs are where the number gets tested.