Build1 publisher3 min readPublished
A RAG stack lived seven hours before a hosted embedding endpoint returned 404
One developer's day-one postmortem is really a dependency inventory: a single third-party call sat under every write and read path, and the fix moved the dependency rather than removing it.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- At 15:53 on March 10, the day after MockEvalio's second commit, the developer shipped a RAG system.
- The shipped system included job description upload with PDF, DOCX and text extraction, a chunker, an embedding service, a vector search layer over pgvector, and a pipeline turning an uploaded job description into five interview questions grounded in it.
- The new part added RagJobDescription and RagJDChunk tables in Postgres, with a 768-dimension vector column.
- Documents were split into 500-800 token chunks with 100 tokens of overlap, and each chunk was embedded and stored.
- A search step pulled back the closest chunks by cosine similarity, cached in Redis, and a generation step combined those chunks with a user's profile and asked an LLM for five questions.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A retrieval pipeline shipped at 15:53 on March 10 and its embedding step stopped working at 22:57 the same evening, when Groq's `nomic-embed-text-v1.5` endpoint returned a 404 [1][6][7]. That is roughly seven hours of uptime [8], and the reason it matters is not the outage but the shape of it: one hosted HTTP call sat underneath both the ingest path and the search path, and the developer, writing on dev.to, says he had not monitored for that specific failure [9].
The stack itself was conventional. Job descriptions came in as PDF, DOCX, or text, went through a chunker, an embedding service, and a vector search layer over pgvector, and came out as five interview questions grounded in the uploaded document [2]. Storage was two Postgres tables, `RagJobDescription` and `RagJDChunk`, with a 768-dimension vector column [3]. Chunks were 500 to 800 tokens with 100 tokens of overlap [4]. Retrieval pulled nearest chunks by cosine similarity with a Redis cache, then generation combined them with a user profile and asked an LLM for the questions [5].
The repair was 17 lines in one file: Groq's endpoint and model swapped for OpenAI's `text-embedding-3-small`, with the 768-dimension output preserved through OpenAI's `dimensions` parameter [10]. That is a small diff because the schema was already pinned to a width that another provider could be talked into producing. It is worth being clear about what it bought, and the author is: the pipeline worked again, and embedding-provider availability is still a dependency, now on whoever is behind the call today [11].
The more useful finding fell out sideways. The old embedding code, when the API key was missing, logged a warning and returned an array of zero vectors, which would have silently corrupted every similarity search built on top of it [12]. The new code throws an explicit exception instead [13]. The author notes this came out of chasing the 404 rather than a deliberate reliability pass [14]. A 404 is loud; a zero vector is not, and a system that will happily persist meaningless embeddings under a missing credential has an availability problem it cannot see.
The behavioural tell arrived six minutes later in the same evening's commits. First OpenAI's Whisper API went in as the primary transcription path with Groq's Whisper as fallback; then a second commit added a self-hosted `faster-whisper` service behind FastAPI in its own Docker container and made that the default, demoting the cloud APIs to fallback [15]. The service's own source comment gives the reason as "minimal cost, no external API calls" [16], and the author treats that as the documented reason rather than necessarily the whole one [17]. Transcription got pulled in-house within minutes of being wired up. Embeddings did not.
What to watch: whether the 768-dimension contract holds if the new provider changes its `dimensions` behaviour, since the schema has no give in it [3][10]; whether the new exception path is exercised before a customer finds it [13]; and whether embeddings follow transcription toward self-hosting or stay a single remote call with no second provider configured [11][15]. The author leaves open why Groq was the first choice for embeddings at all, given that its established strength is inference speed [18], and that question is the one a dependency review would start from.