Build1 publisher3 min readPublished
DuckDB's vss extension removes a database from your RAG stack, then names the price
HNSW indexing runs inside the SQL engine, so there is no second service and no network hop. There is also no index persistence unless you turn on an experimental flag.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- DuckDB ships an official, experimental core extension called vss that adds HNSW-based approximate nearest neighbour indexing to accelerate similarity search over DuckDB's fixed-size ARRAY columns.
- The guide argues that if you are already using DuckDB for analytics or building a lightweight RAG pipeline, there is a good chance you do not need another moving part such as Pinecone, Qdrant, Milvus or pgvector.
- vss implements HNSW (Hierarchical Navigable Small Worlds), described in the guide as the same graph-based ANN algorithm used by most production vector search engines.
- Because the extension is embedded, there is no separate service to run, no network hop and no extra infrastructure to operate; the vector index lives in the same process as the rest of the SQL engine.
- Installation follows DuckDB's usual extension pattern: INSTALL vss; LOAD vss;
Compiled by The EngineerSomething wrong?How this is made
Why it matters
DuckDB ships an official extension, `vss`, that adds HNSW-based approximate nearest neighbour search directly on top of its native fixed-size `ARRAY` type [1]. That collapses a common two-system RAG architecture into one process, and the extension is unusually candid about the conditions under which that collapse is safe [15][17].
The argument, as set out in a practical guide published on dev.to, is that teams already using DuckDB for analytics or a lightweight retrieval pipeline probably do not need Pinecone, Qdrant, Milvus or pgvector as an extra moving part [2]. The mechanism is HNSW, the same graph-based algorithm the guide says most production vector search engines use [3], and because it is embedded there is no separate service to run and no network hop [4].
The workflow is short. `INSTALL vss; LOAD vss;` follows DuckDB's usual extension pattern [5]. You declare the embedding column with its dimensionality fixed at table-creation time, because `ARRAY` is size-constrained, unlike the variable-length `LIST` type [6][7], then build the index with `CREATE INDEX ... USING HNSW` [8]. After that, DuckDB routes any query that orders by a supported distance function against a constant vector and applies a `LIMIT` through the index rather than a full scan [9]. You confirm it with `EXPLAIN` and look for an `HNSW_INDEX_SCAN` node in the plan [10]. The overloaded `min_by(col, arg, n)` aggregate is also index-accelerated and returns the full matched row as a struct [11].
Metric handling is reasonable. The default is `l2sq`, matching `array_distance` [12]; cosine is selected at index-creation time with `WITH (metric = 'cosine')` [13], which the guide calls the natural choice for embeddings from OpenAI or Sentence-Transformers [14]. You can build several indexes on one column with different metrics, though each HNSW index covers exactly one column [15]. `ef_search` can be overridden per connection with `SET hnsw_ef_search`, so accuracy can be traded against latency without a rebuild [16].
Now the part that decides whether this is a real substitution. By default, HNSW indexes can only be created on in-memory databases; persisting one into a disk-backed `.duckdb` file requires setting `hnsw_enable_experimental_persistence` [17][18]. The flag exists because WAL recovery is not yet fully implemented for custom extension indexes, so a crash or kill with uncommitted changes to an HNSW-indexed table can corrupt the index or lose data, and the documentation says this is not recommended for production [19][20]. Read those two facts together and the production-safe configuration is one where the index does not survive a process restart [1]. Recovery from an unexpected shutdown is possible, but it is a manual sequence: start DuckDB separately, load `vss`, and `ATTACH` the database file before WAL replay runs [21]. With persistence on, the whole index is serialised to disk at every checkpoint with no incremental updates, and deserialised back into memory on next access [22].
So the honest boundary is not query semantics, it is durability and working-set size. If your corpus is rebuilt on startup, or small enough to sit in memory and be re-embedded cheaply, the second database is genuinely surplus [2][17]. If you need an index that survives a hard kill, or one large enough that full serialisation per checkpoint hurts, the extension currently tells you so itself [19][22].
Worth watching: whether WAL recovery for custom extension indexes lands and lets the persistence flag lose the word experimental [19], whether checkpoint serialisation becomes incremental [22], and whether `vss` graduates from experimental core extension status at all [1].