Build1 publisher3 min readPublished
Your pipeline's repeat model calls are a cache-key bug, not a quota shortage
A GitLab CI job burned three model calls to produce one answer. Fingerprinting the normalized request, while excluding branch names and build IDs, costs less than begging for more free quota.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- A GitLab CI pipeline that runs a free model over release notes made three model calls for three identical prompts and received the same answer each time.
- The prompt template, model route and input did not change; the job ran again because another branch was merged, and the free quota was consumed.
- The author argues that repeat model spend of this kind is sometimes a quota problem but more often a caching problem.
- Cache keys like release-notes.json are too broad: they collide when the input changes and miss when it does not.
- The recommended cache key is the SHA-256 of the normalized request.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A GitLab CI job that runs a free model over release notes made three model calls on three identical prompts and got the same answer back each time, according to a dev.to post that discloses it was prepared as part of MonkeyCode's product outreach [1][17]. Nothing about the request had changed; the job reran because another branch was merged, and the free quota went with it [2], which makes two of those three calls pure waste [20].
The author's framing is worth stealing even if the vendor pitch is not: this is more often a caching problem than a quota problem [3]. The usual cache key is a filename, and a key like release-notes.json is too broad in both directions, colliding when the input changes and missing when it does not [4]. The replacement is a SHA-256 of the normalized request [5]. In go the prompt template plus rendered input, the model route or identifier, the sampling parameters that deterministically affect output, and the schema version of the parser [6]. Out stay timestamps, build IDs, branch names and random request IDs [7]. That exclusion list is the whole trick: those are the fields that change on every run without changing the question.
The implementation details matter more than the concept. Responses land in .model-cache/<hash>.json [8], and nothing is written unless the response parses, checked by asserting that .answer exists and is a string [12]. That contract check is what keeps a malformed reply from being served for the next day. In .gitlab-ci.yml, the cache key is derived from the prompt file itself, prompts/release-notes/v1.json, under a model-cache prefix, so editing the template invalidates the stored answers without anyone remembering to bump a version [13].
Two gaps in the published script are worth noting before you copy it. The fingerprint JSON it actually builds contains only schema_version, model_route and prompt [9], which means the sampling parameters listed in the design are not in the hash [10]; change temperature and you will serve the old answer. And the entry expires after CACHE_MAX_AGE_HOURS, default 24 [11], which makes the cache time-scoped rather than purely content-addressed, since the same fingerprint can return different results depending on when you ask [21]. That is a defensible choice, not a bug, but it is a different guarantee than "identical request, identical answer".
The single-pipeline version is also fragile by design: GitLab cache is a performance optimization rather than durable storage, and ephemeral runners can start cold [14]. The post's second layer puts the same content-addressed logic behind a small HTTP endpoint, where the client sends the fingerprint and the server returns the stored response or calls the model route, stores the result, and returns it [15]. That is where the vendor arrives, with a free server option to host the shared cache outside the CI worker [16].
The guardrails are conventional and correct: namespace by model route, bump SCHEMA_VERSION when the prompt changes, never cache inputs containing secrets or personal data, and do not cache intentionally non-deterministic tasks [18]. A stale answer is worse than a wasted model call [19].
Watch whether your pipeline reports cache-hit and cache-miss counts at all; without that telemetry you cannot tell a working cache from a silently colliding one. Then check that every parameter your model actually reads is inside the hash.