Build1 publisher3 min readPublished
Codex at OpenAI: stop polishing the prompt, start building the harness
At localhost, OpenAI's Dominik Kundel described a shift from context engineering to agentic delegation. The binding constraint is now whether your repo and tickets can be read by a stranger.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- At localhost, Dominik Kundel walked through how coding with Codex has changed at OpenAI over the past year: autocomplete, then pair programming, and now what he calls agentic delegation.
- Kundel said code has stopped being the bottleneck, and what matters now is whether a team has built the surrounding harness that gives an agent context, a way to validate its own work, and a way to get that work verified.
- Kundel drew a line between this approach and last year's idea of context engineering, hand-crafting the perfect prompt before sending it off, arguing that a colleague who needs everything spelled out isn't much use, and neither is an agent that does.
- Codex needed to work across OpenAI's own large codebase from early on, which meant it had to navigate and understand how things fit together on its own rather than being told which files to open every time.
- Kundel proposed a test: drop a talented new hire into your codebase with nothing but the repository, and see whether they know which tools to use, which conventions apply, and which external dependencies matter.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Speaking at localhost, OpenAI's Dominik Kundel traced a year of change in how his team works with Codex: autocomplete, then pair programming, and now what he calls agentic delegation [1]. His argument, as reported by Render, is that code has stopped being the bottleneck, and that what now decides the result is whether a team has built a surrounding harness that gives the agent context, a way to validate its own work, and a way to get that work verified [2].
That is a direct inversion of last year's advice. Kundel explicitly separated this from context engineering, the practice of hand-crafting the perfect prompt before dispatching it, on the grounds that a colleague who needs everything spelled out is not much use either [3]. The reason is partly historical: Codex had to work across OpenAI's own large codebase early on, so it needed to navigate and understand how things fit together rather than be told which files to open [4]. The test he offers is cheap to run in your own shop: drop a talented new hire into the repository with nothing else, and see whether they can work out which tools to use, which conventions apply, and which external dependencies matter [5].
For most teams some of that lives outside the code entirely, in Linear tickets, Slack threads, or a Google Doc where a decision was made [6]. Kundel's answer is to wire those in, with Notion, Google Drive, Slack, Gmail and his calendar connected to Codex through plugins [7]; he treats tracing a Slack thread as the same skill as tracing a codebase, starting from one message the way you would start from one file [8].
The load-bearing example is a documentation update due to go live at midnight on a Sunday, which he did not want to stay awake for [9]. He prepared the pull request on the Friday, then handed over the rest of the job: watch the relevant Slack channel, get approval from a named teammate and chase a second person if the first went quiet, account for a twenty-minute publishing delay, deploy with enough lead time, confirm the page was live, and post the result to the team [9]. That delay window means the deploy had to start by roughly 11:40 p.m. to hit the deadline [10]. At 12:04 a.m., Codex had verified the page and sent the message [11].
The part operators should sit with is what happened over the weekend. A colleague asked for the preview link; Codex answered in Kundel's place because it was still watching the thread [12]. "That wasn't me," he told the audience. "I wasn't online at that time." [13] Had she said the change looked wrong, he says it would have paused the deployment and worked the fix, because seeing the PR through was part of the assignment [14].
The supporting machinery is mostly context plumbing: App Shots captures application state, not just a screenshot, when you press both Command keys on macOS [15]; Codex can be pulled into a Slack thread, Linear ticket or GitHub conversation as a cloud agent [16]; Memories learns from past interactions such as recurring debugging patterns [17]; and Chronicle watches a person's screen to learn which tools they reach for and in what order [18], which Kundel credits for Codex producing a Google Doc rather than a pull request [19].
What to watch: this is one talk, and the examples are self-reported by the person who ran them. Nothing in the account gives a failure rate, a cost, or what happens when an agent misreads an approval and ships anyway. The transferable part is unglamorous. If the ceiling is set by repo legibility, the work is writing down conventions, making validation runnable without a human, and connecting the systems where decisions actually get made.