Skip to content

Build1 publisher3 min readPublished

Mandi benchmark scores pre-, post- and in-index vector filtering against exact search on 120,000 listings

Open-source benchmark grades pre-filtering, post-filtering and Qdrant's in-index filtering on 120,000 mandi listings against exact brute-force neighbours. Filter placement is the only variable it changes, but the supplied write-up ends before any recall or latency figures.

The Engineer · Build desk

Illustration accompanying Mandi benchmark scores pre-, post- and in-index vector filtering against exact search on 120,000 listings

What happened

  • Each listing's free-text description becomes a 128-dimensional vector, and fields such as transit hours, cold storage, state and tax status become indexed metadata.
  • The post-filtering baseline fetches five times the requested results, discards listings that fail the filter, and retries with more candidates when fewer than ten survive.
  • The repository keeps data generation separate, so a real Agmarknet or eNAM CSV could replace the synthetic listings without changes to the benchmark pipeline.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint With ten results and five-times overfetch, post-filtering fills on its first pass only when about one candidate in five passes the filter, so tight constraints cost extra searches.
  • exposure Teams pre-filtering an HNSW graph can lose true neighbours to pruned links, and the loss only becomes visible when results are graded against exact search, as this run does.
  • capability Because the generator is decoupled, an operator can feed in their own listings and filter mix and get selectivity-specific numbers before committing to one filtering method.

Approximate nearest-neighbour search does not examine every vector [5]. A buyer asking for firm tomatoes for hotel supply may also require transit of eight hours or less, cold storage, Maharashtra and tax exemption [4]. Those hard constraints have to be enforced around a search that skips most of the data by design. The author wrote that the target is "Semantic similarity + hard metadata constraints = usable search result" [14].

Pre-filtering drops the non-matching points and searches what is left. The post's objection is that removing nodes can also remove graph connections the HNSW traversal depends on [8]. The traversal follows those connections. A pruned graph can end the walk before it reaches the true nearest matches [8].

Post-filtering searches first and throws away failures. The post's sketch fetches `top_k=K * overfetch` candidates and keeps the first K that pass [9]. With results capped at ten [10] and an overfetch of five, the baseline pulls 50 candidates and needs ten of them to pass [1]. So the filter has to pass at least one candidate in five, if matching listings are spread evenly through the ranking [1]. When it does not, the post says the system has to retry with more candidates [9]. Five times overfetch is a guess with a retry loop attached. Every retry is another search, and the benchmark's stated question is whether selective filters can be enforced without losing recall or causing latency to explode [6].

The third option puts the filter inside the search call:

`client.query_points(collection_name="mandi_listings", query=query_vector, query_filter=metadata_filter, limit=10)` [10]

The filter becomes part of the search operation instead of an application-level step [10]. Qdrant's HNSW search supports payload filtering, and the benchmark builds payload indexes on `state`, `crop`, `transit_hours` and `moisture_pct` [11]. The call has no overfetch multiplier to tune [10].

The experimental design is careful where it counts. Vectors and HNSW parameters are held identical across methods, and results are scored against exact brute-force neighbours over 60 queries per scenario at four selectivity levels [2]. I think that is the right control. Only the filter's position moves.

Transfer is a separate question. The listings are synthetic but built to be structurally realistic [1]. The generator starts from real crop and mandi reference lists, with Lasalgaon and Pimpalgaon among the onion markets [13]. Descriptions are embedded at 128 dimensions [7]. For a number from this run to carry over, your filters would need to cut your data at selectivities near the four tested here. Your embeddings would need to behave like these short 128-dimensional ones. And your constraint fields would need roughly the same correlation with semantic similarity that the generator built in. The full pipeline is public on GitHub [3]. The supplied text of the post ends inside the generator's reference lists, before any recall or latency table, so which method wins at which selectivity is not in the evidence here [15].

What to watch

  • The recall and latency tables for each of the four selectivity levels, which would show whether in-index filtering holds recall where post-filtering falls back to retries.
  • A rerun of the same pipeline on a real Agmarknet or eNAM export, which would test whether the synthetic data hid correlation between filters and similarity.
  • Publication of the exact HNSW and payload-index settings, which would show whether the single-stage result reflects Qdrant defaults or tuning.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories