Skip to content

Build1 publisher3 min readPublished

A Notion agent that dies after each request, and the debugging error that broke version two

A dev.to write-up describes a one-shot HTTP worker for Notion built on cheap models. The harness held. What broke was three variables, tool surface, model and prompt, treated as one.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A post published on dev.to, headlined 'Building a Disposable Notion Agent on Cheap Models', describes building a one-shot HTTP worker that talks to Notion through MCP.
  • Per the post's TL;DR: version one worked; version two got cheaper and more readable, then failed in a new way.
  • The authors state: 'The harness was fine. The tool surface, the model, and the prompt were not the same problem, and we kept treating them as one.'
  • The requirement was that another service could say 'read this Notion page, write a summary somewhere, stop', with no chat history, no personality that accretes over weeks, no always-on process, and no cost if nobody is calling it.
  • The authors call the shape a 'one-shot agent': one HTTP request, tools for that request, a JSON result, then the instance can go away.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A team writing on dev.to has published the build log for a one-shot HTTP worker that talks to Notion through MCP, and reports that version one worked while version two got cheaper and more readable and then failed in a new way [s1c1][s1c2]. Their diagnosis is the part worth stealing: the harness was fine, and the tool surface, the model and the prompt were not the same problem, but they kept treating them as one [s1c3].

Start with the shape, because it is the opposite of what most agent products sell. The requirement was that another service can say "read this Notion page, write a summary somewhere, stop", with no chat history, no personality accreting over weeks, no always-on process, and no cost when nobody is calling [s1c4]. That is one HTTP request, tools scoped to that request, a JSON result, and then the instance goes away [s1c5]. The stated constraints are as much social as technical: an external caller owns the schedule, the agent is allowed to actually use tools, secrets live in the environment and never in the request body, and idle time is free [s1c7].

The obvious objection is a cron script hitting the Notion API. The authors' answer is that the task changes every call: this week a weekly digest, next week "list in-progress rows and do not write", and they did not want a new Python file per job, but one worker that takes a prompt plus a list of tool servers [s1c8]. They also name the failure mode on the other side, starting from a coding agent with shell, files, memory and a learning loop and shrinking it, where you spend months deleting features you never wanted [s1c9].

Mechanically it is unglamorous. The caller POSTs /run with a system prompt, a user prompt and which tools; the agent starts if needed and dies when idle; Notion sits behind an MCP server holding an integration token; the response returns the result, tools used, tokens and cost [s1c10]. The call is synchronous because a weekly digest can wait two minutes, which removed the need for a job queue in v1 [s1c11]. The agent is not the Notion client, keeping the image thin and the Notion token out of the agent process [s1c12]. Callers name pages in English, so the model searches, picks a title match, and stops if it cannot, refusing to write rather than guess on an ambiguous name [s1c13][s1c6]. The default model was a very cheap DeepSeek flash model behind an OpenRouter-style gateway, on the order of a few cents per million tokens, because nobody wanted a Claude bill on every internal trigger [s1c14].

Auth is where they took the deliberate loss. The self-hosted route was chosen because Notion's hosted connector is built for Claude Desktop and wants OAuth [s1c16]. Rather than per-caller service accounts and short-lived tokens, which they call miserable once the next caller is a script on a laptop, they shipped a shared API key checked at the app and failing closed, accepting the loss of per-caller identity for v1 [s1c17][s1c18][s1c19].

The rule they wrote down was to change model, tool server and prompt before touching the loop, and they say they still almost violated it [s1c15]. That list of three is exactly the list they later admit to debugging as a single variable [1]. When quality wobbled, the temptation was to reach for a "real agent" with memory, skills, a terminal and a personality that improves [s1c20]; the text available to us stops mid-sentence there, so treat the resolution as unreported.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories