Build1 distinct publisher3 min readUpdated
At localhost, OpenAI's Dominik Kundel described a shift from context engineering to agentic delegation. The binding constraint is now whether your repo and tickets can be read by a stranger.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Speaking at localhost, OpenAI's Dominik Kundel traced a year of change in how his team works with Codex: autocomplete, then pair programming, and now what he calls agentic delegation [1]. His argument, as reported by Render, is that code has stopped being the bottleneck, and that what now decides the result is whether a team has built a surrounding harness that gives the agent context, a way to validate its own work, and a way to get that work verified [2].
That is a direct inversion of last year's advice. Kundel explicitly separated this from context engineering, the practice of hand-crafting the perfect prompt before dispatching it, on the grounds that a colleague who needs everything spelled out is not much use either [3]. The reason is partly historical: Codex had to work across OpenAI's own large codebase early on, so it needed to navigate and understand how things fit together rather than be told which files to open [4]. The test he offers is cheap to run in your own shop: drop a talented new hire into the repository with nothing else, and see whether they can work out which tools to use, which conventions apply, and which external dependencies matter [5].
For most teams some of that lives outside the code entirely, in Linear tickets, Slack threads, or a Google Doc where a decision was made [6]. Kundel's answer is to wire those in, with Notion, Google Drive, Slack, Gmail and his calendar connected to Codex through plugins [7]; he treats tracing a Slack thread as the same skill as tracing a codebase, starting from one message the way you would start from one file [8].
The load-bearing example is a documentation update due to go live at midnight on a Sunday, which he did not want to stay awake for [9]. He prepared the pull request on the Friday, then handed over the rest of the job: watch the relevant Slack channel, get approval from a named teammate and chase a second person if the first went quiet, account for a twenty-minute publishing delay, deploy with enough lead time, confirm the page was live, and post the result to the team [9]. That delay window means the deploy had to start by roughly 11:40 p.m. to hit the deadline [10]. At 12:04 a.m., Codex had verified the page and sent the message [11].
The part operators should sit with is what happened over the weekend. A colleague asked for the preview link; Codex answered in Kundel's place because it was still watching the thread [12]. "That wasn't me," he told the audience. "I wasn't online at that time." [13] Had she said the change looked wrong, he says it would have paused the deployment and worked the fix, because seeing the PR through was part of the assignment [14].
The supporting machinery is mostly context plumbing: App Shots captures application state, not just a screenshot, when you press both Command keys on macOS [15]; Codex can be pulled into a Slack thread, Linear ticket or GitHub conversation as a cloud agent [16]; Memories learns from past interactions such as recurring debugging patterns [17]; and Chronicle watches a person's screen to learn which tools they reach for and in what order [18], which Kundel credits for Codex producing a Google Doc rather than a pull request [19].
What to watch: this is one talk, and the examples are self-reported by the person who ran them. Nothing in the account gives a failure rate, a cost, or what happens when an agent misreads an approval and ships anyway. The transferable part is unglamorous. If the ceiling is set by repo legibility, the work is writing down conventions, making validation runnable without a human, and connecting the systems where decisions actually get made.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
At localhost, Dominik Kundel walked through how coding with Codex has changed at OpenAI over the past year: autocomplete, then pair programming, and now what he calls agentic delegation.
Kundel drew a line between this approach and last year's idea of context engineering, hand-crafting the perfect prompt before sending it off, arguing that a colleague who needs everything spelled out isn't much use, and neither is an agent that does.
Codex needed to work across OpenAI's own large codebase from early on, which meant it had to navigate and understand how things fit together on its own rather than being told which files to open every time.
Kundel proposed a test: drop a talented new hire into your codebase with nothing but the repository, and see whether they know which tools to use, which conventions apply, and which external dependencies matter.
Kundel has Notion, Google Drive, Slack, Gmail and his calendar connected to Codex through plugins.
A documentation update needed to go live at midnight on a Sunday; Kundel prepared the pull request with Codex on Friday, then gave it the rest of the job: watch the relevant Slack channel for context, get approval from a specific teammate and track down someone else if that person didn't respond, account for the twenty-minute publishing delay and deploy with enough lead time, confirm the page was live, and post the result to the team.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One vendor-adjacent recap of one speaker's anecdotes
Every claim in the cluster resolves to a single article recapping a single talk by an OpenAI employee about OpenAI's own product. Two demos are described with specific, checkable detail (the 12:04 a.m. confirmation, the twenty-minute publishing delay), but no logs, artifacts, benchmarks, success rates, or second-party accounts are supplied, and the load-bearing generalisations — code is no longer the bottleneck, thread traversal works like file traversal — are asserted rather than measured.
First-party use at the vendor, one production run described
Adoption evidence is real but entirely internal to OpenAI: Codex running against OpenAI's own codebase, Vale linting OpenAI's docs, one staff member's plugin stack, and one delegated midnight deploy. A record-and-replay launch is noted, but no user counts, customer names, team-level rollout, or outside-organisation deployments appear.
Category claim runs ahead of a two-anecdote evidence base
The framing — code is no longer the bottleneck, delegation replaces prompt-crafting — is a general claim about software work supported by two successful anecdotes from the vendor's own staff. The unexercised rejection path, the absence of failure rates, and the silence on permissions and audit for an agent that posts as a human and deploys unattended all leave the narrative overstated relative to the evidence and adoption shown. It is not pure hype: the specific feature and workflow descriptions are concrete and internally consistent.
Vendor employee on stage; publisher's platform adjacent to the closing argument
The speaker works for the company that makes Codex and is describing his own product's capabilities, including a feature launched the same day as the talk. The publisher is a hosting platform whose commercial territory sits directly alongside the article's validation section, which argues for fast test suites and multiple full, independently debuggable environments rather than frontend previews. Both incentives point toward an optimistic reading, which is why the unmentioned failure modes matter.
Low — single source, high incentive alignment, no counter-evidence
Descriptive claims about what Kundel said and what features exist can be held with reasonable confidence because the source is direct and specific. Confidence in the substantive claim — that harness-building is now the binding constraint and delegated agents can be trusted through a deploy — is low, given one source, an interested speaker, an interested publisher, and no independent or contradicting account anywhere in the cluster.
build
Notion's agent stack is live, not slideware, and it only changes one of your decisions1 distinct publisher
security
OpenAI's Computer History writes a plaintext log of the workday. Decide before staff opt in.1 distinct publisher
product
Record, don't prompt: two labs converge on demonstration as the agent interface1 distinct publisher
invest
DeepSeek V4 Flash costs a tenth as much and passes 53.8% of agent tasks1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 16, 2026