Build1 publisher3 min readPublished
Your first MCP workflow should be a draft queue, not an agent with keys to the inbox
A dev.to post makes the operational case: a wide tool surface turns every request into a routing problem, and afterwards nobody can say which tool got called or what it changed.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The Model Context Protocol gives apps a standard way to expose tools to language models; a tool has a name, description, and input schema, and can query a database, call an API, or run a computation.
- MCP standardisation does not decide which tools an agent should see, or which calls may change state; those are application decisions.
- The tempting demo wires the agent to CRM, inbox, calendar, WordPress, analytics and payments; it looks powerful but creates a large surface for wrong tool choices, accidental writes and duplicate sends.
- A wide tool surface leaves questions nobody can answer later: which tool did the model call, what information did it use, what would it have changed if nobody stepped in, and can we replay the decision without dumping the whole customer record into a log.
- For a small team the safer start is a draft queue: the agent researches and prepares a proposed action, a person approves it, and only then does a narrow workflow perform the side effect.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A dev.to post published under the handle sphillips1337 argues that the first useful Model Context Protocol project for a small business should be a draft queue with a human approval step rather than an autonomous agent: the model researches and prepares a proposed action, a person approves it, and only then does a narrow workflow perform the side effect [5]. The author's framing is that most small teams do not need autonomy, they need the next customer reply drafted and the right product notes found without anything odd going out overnight [20].
The protocol itself is narrower than the marketing suggests. MCP standardises how an application exposes a tool to a language model, with a name, a description and an input schema, behind which sits a database query, an API call or a computation [1]. It does not decide which tools an agent should see or which of those calls may change state; those stay application decisions [2]. As the post puts it, the protocol can make integrations interoperable but cannot make a sloppy permission model safe [19].
That is where the routing problem starts. The tempting demo wires the agent to CRM, inbox, calendar, WordPress, analytics and payments, which the author says creates a large surface for wrong tool choices, accidental writes and duplicate sends [3]. It also leaves four questions with no cheap answer afterwards: which tool did the model call, what information did it use, what would it have changed if nobody stepped in, and can the decision be replayed without dumping the whole customer record into a log [4].
The sorting rule offered is blunt. Every tool is either read or prepare (search the knowledge base, look up an order, draft a reply) or side effect (send email, publish, refund, update CRM, delete), and only the first category is a reasonable pilot surface; the second sits behind an explicit approval boundary until the workflow has earned trust [6]. In the worked example, a five-person agency handling website enquiries needs four tools: search_services, find_faq, lookup_enquiry and create_reply_draft [7]. On the author's own taxonomy that is three reads and one prepare, and no side-effect tool at all [21]. Absent by design: no send-email tool, no unrestricted filesystem access, no read-every-customer, no inventing prices from a private spreadsheet [8]. The worker takes an enquiry ID rather than a pasted customer dump, drafts a reply with links to the source notes it used, stores it with status needs_review, notifies a reviewer, and a separate deterministic workflow sends the exact approved text [9]. Approval is a state transition, not a polite line in the system prompt [10].
Before drafts reach production, the post recommends shadow mode: the agent runs on real or representative requests, cannot create even a draft, and writes proposed tool calls and outputs to a review log a human compares against their own answer [11]. Only after acceptable shadow results should it create needs_review drafts, with automatic side effects later still [14]. The approval record should be durable rather than buried in a chat transcript, and the fields that matter are action, scope, source references and state [15][16]. Reviewer edits get recorded as both feedback and audit trail [17], and the send step is a narrow integration with boring validation that accepts a queue ID [18].
What to watch is the two-week evidence the author proposes and does not report: how often the agent picked the right source, how often drafts needed a factual correction, which requests were ambiguous, whether it asked for data it should not have needed, and how many proposed actions would have been unsafe if executed [12]. The most useful signal is the edit rate, because repeated corrections usually indicate a documentation problem rather than a model problem [13]. Teams keeping that log should also check that no sent item ever carries a null approved_by [15][16].