Readyset's compiler rewrites supported SQL for incremental caching
Co-founders Alana Marzoev and Jon Gjengset turned their MIT Noria research into Readyset. In an April 28th engineering article, Vassili Zarouba explains how its compiler rewrites supported SQL for incremental caching.
By Ryan Merket · Published
Primary source: Readyset
Why it matters
Readyset's route to database-cache adoption runs through SQL compatibility: each semantics-preserving rewrite can bring more existing application queries within reach, while unsupported patterns still bypass the cache.

Readyset's query compiler turns supported SQL into continuously maintained caches, a product rooted in co-founders Alana Marzoev and Jon Gjengset (@jonhoo)'s MIT research. In an engineering article published on April 28th, Readyset engineer Vassili Zarouba detailed how the system rewrites queries into forms its dataflow engine can keep current as the underlying database changes.
A conventional database runs a query when asked. Readyset can build a dataflow graph for a cache and update its result as rows are inserted, changed or deleted. A read from that cache becomes a lookup into maintained results instead of a fresh execution of the full query. That approach is the commercial extension of Noria, the streaming dataflow research project Marzoev and Gjengset worked on at MIT. Readyset's own early description of the product traces Readyset's origins to that research and its founders' effort to automate cache maintenance.
Rewriting SQL around the engine
SQL lets developers express the same task in many different forms, including nested queries, common table expressions and joins. Readyset's dataflow engine has more specific structural needs: it represents a cached result as a graph of operators that process changes over time. Zarouba's post says the engine favors two-input joins connected by equality predicates and cannot execute a correlated subquery once for every row in an outer query. The rewrite pipeline converts compatible SQL shapes into a form the engine can compile.
The article groups those transformations into three stages. Normalization resolves tables and columns, expands SELECT * and converts syntax such as JOIN ... USING into explicit join conditions. Deep rewrites handle more complex changes, including turning some subqueries into joins and flattening derived tables when doing so preserves the query's meaning. Cleanup strips redundant clauses and parameterizes literals so structurally similar queries can share a dataflow graph.

Readyset aims to fit between an application and an existing MySQL or PostgreSQL database, offering an alternative to building and maintaining application-level caches. But the query must fit the cache engine's supported shapes to receive this kind of incremental maintenance. A SQL statement can be valid in the upstream database and still fail to qualify for a deep cache. Readyset says unsupported queries continue through to the database rather than being cached.
Compatibility is an engineering roadmap
The article's limitations are a snapshot, not a permanent product boundary. Readyset's August 20th follow-up by Zarouba describes new rewrites for subqueries in HAVING, ORDER BY, and both inner and left-outer join conditions. It also explains why these changes require careful handling: an apparently simple rewrite can change how a left join preserves rows or how NULL values affect a result. In cases the system cannot safely transform, the query is sent to the upstream database without a cache.
Current Readyset query-support documentation reports tests against PostgreSQL 17 and MySQL 8.4 as of October 5th, 2026. It lists cache support by specific query shape and distinguishes deep caches, which are incrementally maintained, from shallow caches, which refresh on a time-to-live schedule. The documentation also provides a way to check an individual query with EXPLAIN CREATE CACHE or inspect queries that are being proxied. Operators can use those details to check whether their workload can be cached and which cache mode applies.
Marzoev's founding thesis was that developers should not have to build and maintain bespoke cache-invalidation systems to scale database reads. The current pipeline addresses the engineering work behind that thesis: expanding the set of SQL patterns the product can safely support without changing what those queries mean. Gjengset's and Marzoev's research project supplied the technical foundation; Zarouba's engineering posts describe the ongoing task of making the system accept more of the SQL developers already write.
Readyset competes for teams already running relational databases and weighing custom caches, replicas or query changes. The more existing queries it can safely maintain, the less application-specific adaptation customers need before trying it. Readyset raised a $24 million Series A led by Index Ventures in April 2022, with participation from Amplify Partners, according to TechCrunch's coverage. That financing is historical context, not evidence of current traction or performance.
The April pipeline article explains the machinery; the August follow-up shows that the supported boundary can move. Neither post independently benchmarks the system's performance in customer workloads, so the strongest evidence for any deployment remains whether the application's own queries pass Readyset's support checks and deliver the expected result.