Skip to content

Build1 publisher3 min readPublished

Turbopuffer v3 demotes the ANN index to one secondary index among several

Turbopuffer is rewriting its serverless database as v3, moving the ANN index out of the core of storage and making it one of several secondary indexes. A dev.to explainer of the September 30 post says the vector-first layout held back GROUP BY and aggregation queries.

The Engineer · Build desk

Illustration accompanying Turbopuffer v3 demotes the ANN index to one secondary index among several

What happened

  • Each namespace lives in object storage as a set of immutable files, and every write adds a new file instead of changing an existing one.
  • Queries check a RAM cache, then an NVMe cache, and read object storage only after missing both, promoting what they fetch to the faster tiers.
  • A direct object-storage read takes tens to hundreds of milliseconds, too slow on its own to serve interactive search.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint On data laid out by vector cluster, an aggregation over an ordinary attribute cannot skip any files, so on a cold namespace it pays object-storage read latency across the whole dataset.
  • decision Teams that run a vector store beside a second database for filters and rollups now have a reason to test one system for both, once Turbopuffer publishes v3's query semantics and numbers.
  • precedent A vector-database vendor demoting its own vector index treats similarity search as one feature of a general store. Buyers can now ask competing vendors how their engines run GROUP BY.

Turbopuffer titled its September 30 post "RIP, vector database" [1]. That is a bold title for a company still described as a serverless vector database in the dev.to explainer of the post [1]. Underneath the title, the storage was already built like a general database.

Writes never update in place. Each write to a namespace adds a new immutable file to object storage, the way Git adds commits instead of rewriting history [5]. Object storage bills for gigabytes held and for read and write operations. It charges nothing for an idle server [6]. A namespace nobody queries therefore costs close to nothing [11]. That matters for RAG and agent applications, which can give each user, document or conversation its own isolated namespace [12]. Keeping thousands of small namespaces live in memory is not viable on price. Serving them on demand from object storage is [12].

The cost of that design is latency. A direct object-storage read takes tens to hundreds of milliseconds [7]. A query works down three tiers:

1. RAM, the fastest and most limited tier. 2. NVMe, slower than RAM but cheap enough to hold whole namespaces that do not fit in memory [13]. 3. Object storage, the source of truth, read only after a miss in both caches. The fetched data is promoted to the tiers above [4][8].

This is sound engineering. The files never change, so a cached copy cannot go stale. The caches need promotion and eviction logic, and no invalidation protocol.

The vector index is SPANN and SPFresh, which group vectors into a hierarchical tree of clusters where HNSW builds a graph [9]. In my view that choice suits object storage. A query can fetch a few whole clusters near its vector, while a graph walk turns each hop into another dependent read.

According to the explainer, the vector-primary design held back GROUP BY and aggregations over the same data [3]. It does not explain the cause. I think it is the layout. If data is grouped by vector cluster, rows that share an attribute value are spread across clusters, and an aggregation over that attribute has to touch all of them. On a cold namespace each touch can be an object-storage read at tens of milliseconds or more, billed per operation [6][7]. v3 makes the ANN index one of several secondary indexes and takes it out of the center of the system [2]. The explainer does not say what the primary organization becomes, which other indexes v3 adds, or how aggregation latency changes.

The broader case, that retrieval stacks are turning into general-purpose databases on object storage, rests on one vendor's rewrite, reported secondhand. One detail points the same way. Linear's sync engine is one of three named workloads on the architecture, beside Cursor and Notion [10]. If that workload is mostly reads and writes of records by key, customers were already using the system for more than similarity search.

What to watch

  • Turbopuffer's own v3 documentation describing the new primary data layout and listing the other secondary indexes.
  • Published latency and per-query cost figures for GROUP BY and aggregations on cold namespaces served from object storage.
  • Whether and how existing Cursor, Notion and Linear namespaces are migrated to v3.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories