Build1 publisher2 min readPublished
Lakebase swaps Postgres copy-and-replay restores for a branch at a timestamp
Databricks says Lakebase restores Postgres in seconds even at 100 TB by creating a branch at a timestamp. The figure times branch creation, so it matches a finished RDS restore only if that branch serves production reads at once.
The Engineer · Build desk

What happened
- On RDS, a point-in-time restore starts by provisioning a new instance at least as large as the primary, and large EBS volumes take longer to come up.
- RDS then marks the instance available while the S3 snapshot is still hydrating, and fetches any block a query touches from S3 on demand.
- Failing over to a healthy replica does not help when a dropped table or bad write has already reached the standby, so those cases still need a restore.
- Databricks says that unless a database is small, point-in-time recovery on that path is almost always a multi-hour operation.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- capability If branch creation stays a metadata operation, restore time stops tracking database size, and the largest databases are no longer the slowest to recover.
- cost Every RDS restore stands up a second instance at least as large as the primary, so the biggest databases pay the most for each recovery in both money and hours.
- constraint Lakebase's seconds figure counts as recovery only once a fresh branch serves production reads at normal latency, the same bar Databricks sets for RDS.
- decision For teams staying on RDS, snapshot age controls replay time, since an hour-old snapshot leaves far less WAL to replay than last night's.
Lakebase splits compute from durable storage and connects the two through the WAL [4]. The compute node does the familiar Postgres work. It runs SQL, plans queries, applies MVCC, manages locks and generates WAL [4]. Database history is already in object storage, kept in a form Databricks says can be referenced instantly [3]. To restore to a moment T, Lakebase creates a branch at T [1]. According to Databricks, the branch is a metadata operation and no data is copied onto a new disk [1][3].
Postgres point-in-time recovery has two ingredients: a base backup and the WAL archived after it. On RDS the base backup is a snapshot in S3 [13]. Databricks says that path gets slower and more expensive as the database grows [14]. Branch creation moves no pages. Databricks' claim of seconds at 100 TB rests on that [2].
I think the design is sound. The costly step, getting history into object storage in a form you can reference, happens before anything goes wrong [3]. The restore path has nothing left to copy.
The number needs one check before it transfers. Databricks holds RDS to a strict test. By its account, a restore is only done when the data you actually need is on the volume, and on-demand S3 fetches have latency that is fine for an internal checkup but not for production [8]. A Lakebase branch also reads from durable storage that sits apart from compute [3][4]. So a branch created in seconds and a database recovered in seconds are different endpoints [1]. The part of the post describing restores does not report query latency on a fresh branch.
For the 100 TB figure to hold for a given workload, the storage layer has to answer a new branch's first reads within that application's latency budget. The restore clock also has to stop where Databricks stops it for RDS: when production traffic runs normally [1].
Databricks wrote that the restore is "so simple that an agent can do it" [11]. An agent could run an RDS restore too, and it would wait through the same multi-hour window [10]. On the RDS side, Databricks backs its case with a survey of 50 developers running production Postgres at 1 TB or more [12].
What to watch
- Measured first-query latency on a freshly created Lakebase branch, compared with a fully hydrated RDS volume of the same size.
- Publication of the results and method from Databricks' survey of 50 developers running 1TB+ production Postgres.
- Pricing for the compute attached to a restore branch, compared with the full-size instance an RDS restore requires.