Build1 publisher3 min readPublished
DynamoDB vector indexes remove the second datastore, and the GSI permutation trap with it
A walkthrough of bolting semantic search onto an existing serverless recipes API shows the sync pipeline is now optional. The infrastructure-as-code story has not caught up yet.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- A dev.to post in the AWS Builders space describes taking an existing DynamoDB table and adding vector search to it, so users can search recipes with natural language queries and find results based on meaning rather than exact keyword matches.
- The author says the problem with adding search is not the implementation but everything that comes with it: extra components to manage, more failure points, and the constant challenge of keeping data in sync.
- The existing API uses a serverless setup with SAM for infrastructure, API Gateway in front, Lambda functions behind, and DynamoDB for storage, exposing plain CRUD operations to create, read, update, delete and list recipes.
- Adding filters in DynamoDB requires a Global Secondary Index for every permutation, which the author states is really not scalable.
- According to the author, the only option until now was to have a data pipeline that indexed the data separately and provided search, and that is no longer the case with vector search in DynamoDB, where embeddings are stored alongside the data and searched directly.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A post on dev.to walks through taking an existing DynamoDB-backed recipes API and adding semantic search to it directly, with no separate search service [s1c1]. The interesting part is not the embeddings, it is what gets deleted from the architecture: the second datastore and the pipeline that kept it in sync [s1c2].
The author's starting stack is unremarkable in the best way: SAM for infrastructure, API Gateway in front, Lambda behind, DynamoDB for storage, and plain CRUD operations to create, read, update, delete and list recipes [s1c3]. The failure mode is familiar too. Adding filters in DynamoDB means adding a Global Secondary Index for every permutation, which the author describes as not scalable, so the only previous option was a separate data pipeline that indexed the data elsewhere and served search from there [s1c4][s1c5]. That is the honest reason search so often lives outside the primary database. It is rarely a query-language problem. It is a fan-out-of-indexes problem, and the escape hatch has historically been a copy of your data somewhere else.
The mechanics are deliberately boring. You store the embedding as an attribute on the item, create a vector index over that attribute, and query it with a dedicated similarity API, which the author compares to creating a GSI and calling it with Query [s1c6]. The embedding model is Amazon Bedrock's Titan Text Embeddings V2, producing 1024-dimension vectors returned normalized, which pairs with cosine similarity [s1c7]. The index was created with cosine distance, 1024 dimensions to match the model output, and an inline filter on cuisine; the author likens the inline filter to a partition key in a regular index, except it is not required [s1c8][s1c9].
The design decision that actually determines search quality is the input text. Structured recipe fields are flattened into a single string covering name, description, cuisine, dietary tags, ingredients and timings, so one query can match on any of them at once [s1c10]. That is six field groups collapsed into one vector [s1c11], which is the direct substitute for the GSI-per-filter sprawl.
Two consequences worth pricing in. First, the embedding is generated inline on write, which is what makes every item searchable the moment it is written [s1c12]; it also means a model invocation now sits on the write path, so write latency and write failures inherit Bedrock's behaviour. Second, vector indexes are not yet supported by CloudFormation, so the author could not declare the index in the SAM template and instead added a post-deployment script that creates it with UpdateTable if it does not already exist [s1c13]. A feature that lives outside your template is a feature that drifts, and a create-if-not-exists script is a workaround with a shelf life.
Watch for CloudFormation and SAM support for vector indexes, which is the gate on this being a first-class part of a deployment rather than a bolt-on step [s1c13]. Watch the inline filter choice as well: it is declared when the index is created, alongside the attribute definition [s1c8][s1c9], so picking the wrong prefilter dimension is a schema commitment, not a query-time decision. And watch what happens to teams whose embedding text definition changes after launch, because the field list is application code [s1c10] while the stored vectors are data.