Skip to content

Build1 publisher3 min readPublished

Grill the plan first: a narrow interrogation step for agent-assisted builds

A dev.to post describes grill-me, a manually invoked skill that interviews you one question at a time and inspects the repo instead of asking you to recall it, before code exists.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A dev.to post titled "Use grill-me to Pressure-Test an AI Implementation Plan Before Code" describes grill-me as a manually invoked skill that interviews the user about a plan or design until the important decision tree is resolved.
  • Many software mistakes begin as decisions that nobody explicitly made: a feature request sounds clear enough, an AI coding agent begins implementation, and the details get settled by whichever model output appears first.
  • Teams later discover that "add roles", "cache this endpoint" or "support collaboration" contained several linked product, data, security and rollout choices.
  • The skill asks one question at a time, supplies a recommended answer, and waits for feedback before continuing.
  • If an answer can be discovered by inspecting the codebase, the agent should investigate rather than ask the user to recreate repository facts from memory.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A post on dev.to describes grill-me, a manually invoked agent skill whose only job is to interview you about a plan or design until the important decision tree is resolved [1]. It matters because the expensive failure in agent-assisted development is rarely bad syntax: a request sounds clear enough, the agent starts implementing, and the details get settled by whichever model output appears first [2].

The diagnosis is the useful part. Phrases like "add roles", "cache this endpoint" and "support collaboration" each carry several linked product, data, security and rollout choices, and teams find out later [3]. Nobody made those decisions; they were absorbed by the first draft of the code.

The mechanics are deliberately small. The skill asks one question at a time, supplies a recommended answer, and waits for feedback before continuing [4]. The rule that does the most work is the one about evidence: if an answer can be discovered by inspecting the codebase, the agent is supposed to investigate rather than ask you to recreate repository facts from memory [5]. Before asking whether an endpoint follows REST or RPC conventions, it should read the existing endpoints [18]. Anyone who has watched an assistant ask "what is your current auth model?" while sitting inside the repository will recognise the saving.

Sequencing is the second design choice. According to the post, a long questionnaire feels efficient but usually is not, because it asks about details before the premise is fixed and the answers end up inconsistent; a sequential interview lets dependencies resolve in order [9]. The worked example is organisation-level roles: first, whether roles are global or scoped per organisation, which determines whether membership is a separate domain entity; then whether permissions are static bundles or configurable; only then API shape, migration, admin UI and audit [10].

The question set clusters into five areas: scope, data and state, authorization, failure modes, and rollout [19]. Scope asks what behaviour changes, who can trigger it, who can observe it, and what is explicitly out, which is what stops a small feature becoming a platform redesign [13]. State asks which transitions are legal, and whether resending an invitation mints a new token, extends expiry, or reuses the record, which are product rules with storage consequences [14]. Authorization asks who may act, at which boundary, and what a denial reveals, on the argument that it is not a closing middleware detail [15]. Failure modes cover the email already attached to a member, delivery failures, two admins acting at once, and client retries [16]. Rollout may well conclude that nothing special is needed, but consciously [17].

The post is explicit about what this is not: not a general implementation workflow, and not a replacement for a specification, issue breakdown, tests, or code review [6]. It is a pressure test for a direction not yet settled enough to build, on the view that teams often need an agent to push back rather than agree, since a conventional assistant leans toward accepting the first plausible framing [7][8].

Two things to watch. The skill is manually invoked, so it only helps operators who already suspect their plan is soft, and the sample prompt front-loads the outcome, affected users, constraints, existing artifacts and the specific decision at issue [11][12]. And the post reports a practice, not a result: there are no measured outcomes in it [20]. The test worth running is whether the recorded settled decisions and open risks [12] survive into the implementation the agent then writes.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories