Skip to content

Build2 publishers3 min readPublished

Databricks' Lakebase Search moves Postgres vector indexes from RAM into object storage

Databricks made Lakebase Search generally available, pairing BM25 with a Postgres vector index it says runs 4 times cheaper than pgvector at 100M vectors. The vector index lives in object storage behind a cache, so large agent corpora no longer need RAM sized to the whole index.

The Engineer · Build desk

Illustration accompanying Databricks' Lakebase Search moves Postgres vector indexes from RAM into object storage

What happened

  • Lakebase Search ships as two Postgres extensions, lakebase_vector for approximate nearest-neighbor search and lakebase_text for BM25 full-text search, on AWS and Azure.
  • Existing pgvector users switch by adding a lakebase_ann index while keeping the same vector types, distance operators and query syntax, with no data migration.
  • lakebase_text adds a lakebase_bm25 index that keeps standard Postgres tsvector types and operators while adding corpus-wide BM25 ranking and top-K pushdown.
  • Databricks says customer Conexiom runs BM25 hybrid search on over 100 million rows using half the compute of its previous pgvector setup.
  • Neon pairs the extensions with TypeScript Functions for ingestion, an OpenAI-compatible embeddings endpoint in AI Gateway, and a trigger that indexes files uploaded to Object Storage.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Trying lakebase_vector on an existing pgvector table costs an index build and a recall check on real queries; the schema and application code stay as they are.
  • cost Agent databases that idle between bursts no longer have to keep always-on compute sized to the whole vector index, since the index persists in object storage through suspends.
  • constraint The fourfold cost claim is a 100-million-vector result against an unnamed pgvector vendor, so teams whose corpora fit in RAM have little evidence it applies to them.
  • capability An agent's retrieval tool can pull semantic and exact-term matches from the same Postgres that holds operational data, without an ETL pipeline into a separate search engine.

Databricks starts its case with pgvector's memory model. By its count, a 768-dimension float32 vector needs about 3.3 KB once HNSW graph links and Postgres overhead are included, so 100 million rows need roughly 330 GB of RAM to keep the index resident [13]. When the index spills to disk, Databricks says, queries slow by 10x to 50x [14]. HNSW has no global rebalancing, so restoring search quality takes a full REINDEX that locks the table and blocks writes [16]. Each query also runs in one backend process, and the index scan is never parallelized [17]. Databricks is grading an index it now competes with, though pgvector is also the most installed extension on its own platform [20].

The design follows from where Lakebase keeps data. Durable data sits in object storage, with RAM and local NVMe acting as caches in front of it. On that layout, Databricks writes, HNSW search means a series of random object store reads [18]. lakebase_vector uses hierarchical IVF to narrow each search to a small set of clusters, and RaBitQ quantization to compress vectors so more of the index fits in the cache layers [3]. The durable index stays in object storage [2]. I think IVF is the right choice for a database built this way. A cache in front of object storage handles a few cluster reads better than a chain of random hops through a graph.

Scale to zero is the piece of engineering I would credit first. When compute suspends, the index stays in object storage, and returning traffic reconnects to the same index instead of rebuilding it [4]. By Databricks' own figure, a pgvector index build that spills to disk takes almost 50 hours on a standard cloud instance [15].

The headline benchmark is a VectorDBBench run that Neon says used LAION-100M [10]. Databricks reports twice the throughput of the next best system, and sets the fourfold cost saving against "a cloud Postgres vendor using pgvector", before any autoscaling savings [8]. It reports P99 latency of 71 milliseconds at 97% recall [9]. That leaves about 3 in every 100 true nearest neighbors unreturned [1].

For the result to transfer, a corpus has to be large enough that an HNSW index would spill out of RAM. The recall target has to sit near 97%. Queries also have to keep hitting a working set small enough to stay cached. Databricks' per-vector figure puts a 10-million-row index near 33 GB [2], and at that size I'd expect pgvector's spill penalty to matter much less.

Hybrid search is plain SQL. An application combines vector and keyword results in one query [7]. Neon's example wraps that query in a Function that the agent calls as one tool, and the call returns ranked passages with their sources [21]. The Neon post does not say how the two rankings are merged into a single order.

Jordan Voves, an AI/ML architect at Conexiom, said in the Databricks post: "Lakebase Search gives us a whole new level of scalability over pgvector, and unlocks BM25 in the same serverless database. We use Lakebase to connect data to our agents at scale." [12]

What to watch

  • An independent VectorDBBench run that names the pgvector vendor and instance sizes behind Databricks' 4x cost figure.
  • Documentation of how lakebase_vector and lakebase_text scores are fused into one ranking in hybrid queries.
  • Recall and latency figures for lakebase_vector at corpus sizes well under 100 million vectors, where pgvector's index fits in memory.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories