Skip to content

Build1 publisher3 min readPublished

S3 Vectors pre-filtering reaches existing indexes only after they move to ENHANCED mode

AWS says metadata pre-filtering in S3 Vectors returns up to 5x more matching vectors on highly selective filters, at no additional cost. Existing indexes keep CLASSIC search until updated, so multi-tenant RAG stores get the gain only after an index-mode switch.

The Engineer · Build desk

Illustration accompanying S3 Vectors pre-filtering reaches existing indexes only after they move to ENHANCED mode

What happened

  • On ENHANCED indexes, S3 Vectors resolves the filter first and searches only matching vectors, while CLASSIC checks each candidate against the filter during the search.
  • According to AWS, the feature requires no re-ingestion of stored vectors and no change to query code.
  • AWS tells users to grant IAM permissions for the new actions before they start.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Each existing index is its own migration: query code stays the same, but results under selective filters change once the owner flips the mode.
  • constraint The 2 KB ceiling caps how much scoping data each vector can hold, so per-document access lists have to stay short or be replaced by a coarser tenant key.
  • capability Prefix matching lets one expression scope a query to a folder or matter-number subtree, where listing every ID would eat into the 100-constraint budget.

"Most applications never search a whole index," AWS wrote in the launch post [11]. S3 Vectors already took metadata filters. On a CLASSIC index it runs the vector search and the filter check together, validating each candidate against the filter as the search proceeds [4]. The candidates are drawn from the full index [9]. The filter only decides which of them survive.

AWS's own worked example shows where that breaks. A support knowledge base holds 8 million tickets, and one customer owns 400 of them [9]. That customer is one ticket in 20,000, or 0.005% of the index [1]. Under CLASSIC, AWS says, the query drew its candidates from all 8 million and returned fewer of that customer's matching tickets [9]. At that ratio I'd expect most of the near neighbours the search visits to belong to other customers and get discarded. On an ENHANCED index, S3 Vectors resolves customer_id first, and the similarity search then covers all 400 [9][4].

The headline figure is "up to 5x more of the matching vectors" on highly selective filters [2]. That is a ceiling from AWS's queries. For it to hold on your workload, two things have to be true. Your filters have to be narrow, like the single client inside a firm-wide archive that AWS says is where recall improves most [10]. And your CLASSIC results have to have been short of matches in the first place. A category filter that matches a large slice of the index already passes most candidates, so I'd expect a smaller gain there. AWS did not publish the dataset, the top-k, or latency for either mode behind the 5x number.

I think resolving the filter first is the right default for tenant-scoped retrieval. Putting it behind a per-index mode also means nobody's result sets change until the index owner chooses it [4][5]. AWS says there is no additional cost, no re-ingestion and no change to queries [3]. Two steps still sit with the operator:

1. Grant IAM permissions for the new actions [6]. 2. Update each existing index, because existing indexes use CLASSIC until updated [5].

The filter has hard limits. Each vector carries at most 2 KB of filterable metadata, and a query accepts at most 100 filter constraints [7]. The walkthrough's PutVectors call attaches four fields: tenant_id, category, created_date and active [12]. That fits with room to spare. A design that stores a per-document access list in metadata will reach 2 KB much sooner. The new $startsWith operator matches prefixes of paths, URLs and hierarchical keys [8]. AWS's legal example uses it to narrow one client's documents by matter number, folder path or document ID prefix [10].

What to watch

  • Whether AWS publishes the dataset, top-k and latency behind the up-to-5x recall figure.
  • Whether newly created S3 Vectors indexes default to ENHANCED or need the mode set explicitly at creation.
  • Query latency on ENHANCED indexes for broad filters that match a large share of the index.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories