Skip to content

Build1 publisherNot yet confirmed elsewhere2 min readPublished

Readyset widens its cache by rewriting more SQL into shapes its dataflow engine maintains

Readyset's compiler now rewrites subqueries in HAVING, ORDER BY and inner and outer join conditions so its dataflow cache can keep them current. Queries it cannot safely rewrite go to the database uncached, so SQL coverage sets how much read load it takes.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Readyset's cache engine is the commercial extension of Noria, the MIT streaming dataflow project of co-founders Alana Marzoev and Jon Gjengset.
  • Engineer Vassili Zarouba's April 28 post splits the compiler into three stages: normalization, deep rewrites and a cleanup pass.
  • His August follow-up warns that a simple-looking rewrite can change which rows a left join keeps or how NULL values affect a result.
  • Readyset's query-support docs list cache support by query shape, tested against PostgreSQL 17 and MySQL 8.4 as of October 5, 2026.
  • Operators can check a single query with EXPLAIN CREATE CACHE or inspect the queries Readyset is currently proxying to the database.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint An existing app gains only on queries whose shapes compile, so a workload heavy in unsupported patterns keeps hitting the database however fast the cached path is.
  • cost Queries that end up in shallow caches refresh on a timer, so their results can lag the database between refreshes, which is the staleness incremental maintenance is meant to remove.
  • decision Sizing Readyset for an app starts with sorting its query log into deep, shallow and proxied, since only the first group gets incrementally maintained results.
  • precedent Each rewrite release can move proxied queries into the cache, so a workload's fit is a dated result to recheck against the current support list.

A conventional database executes a query when it arrives. Readyset builds a dataflow graph for each cache and updates the stored result as upstream rows are inserted, changed or deleted [3]. A read is then a lookup into a result that already exists [3]. Readyset sits between the application and an existing MySQL or PostgreSQL database [19].

The graph sets the limits. It is a set of operators that process changes over time [18]. According to Zarouba's April 28 post [2], the engine favors two-input joins on equality predicates and cannot execute a correlated subquery once for every row of an outer query [5]. SQL lets developers write the same task as a nested query, a common table expression or a join [12]. The engine takes only some of those forms. The compiler's job is to translate the others without changing their results [1].

It does that in order:

1. Normalization resolves tables and columns, expands SELECT * and rewrites JOIN ... USING as explicit join conditions [6]. 2. Deep rewrites turn some subqueries into joins and flatten derived tables where the meaning holds [21]. 3. Cleanup strips redundant clauses and replaces literals with parameters so structurally similar queries share one dataflow graph [7].

The third step is easy to overlook. Without it, each distinct literal value would describe its own query shape and its own graph [7].

The supported set moves between posts. Zarouba's August 20 follow-up [9] adds subquery rewrites in four positions [22], each a place where a careless rewrite returns a different answer [10]. I think the fallback is the best decision in the design. When the compiler cannot transform a query safely, the query goes to the upstream database without a cache [11]. A declined rewrite costs the application the same database query it was already running.

Speed is the wrong first question for this product. Any latency figure for a Readyset read describes queries that compiled [1]. For that figure to transfer to an existing app, the app's hot queries have to compile too. The share of read load Readyset absorbs is the share of query executions whose shapes the compiler accepts. A query can be valid upstream and still miss a deep cache [8].

Marzoev's founding thesis was that developers should not have to build and maintain their own cache-invalidation systems to scale reads [17]. The invalidation logic still has to exist somewhere, and Readyset's bet is that one compiler team writes it once [17]. The pitch goes to teams already weighing custom caches, read replicas or changes to their queries [16].

What to watch

  • Whether the documented shape list grows to cover correlated subquery forms beyond the four positions in the August post.
  • Customer offload or hit-rate figures measured with the proxied-query view on production workloads.
  • Support testing against PostgreSQL and MySQL releases newer than 17 and 8.4.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories