Skip to content

Build1 publisher3 min readPublished

Codex learns to click: the coding agent stops typing patches and starts operating the machine

OpenAI's Codex update adds computer use on Windows and remote control, per a dev.to write-up. That changes what an engineering team hands off, and how much a host machine can be trusted.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A May 29 Codex update from OpenAI added support for computer use on Windows in the Codex app for eligible users, so it can see, click and type in Windows applications while testing and refining software.
  • The same release expands remote control, letting a user steer work from ChatGPT on mobile or Codex on Mac while the Windows machine remains the host for the project files, shell, app server and local context.
  • The author argues the most interesting thing about AI coding agents is not that they can write code but that they are starting to operate computers.
  • The author contrasts a code generator, which lives inside a text box, waits for a prompt, returns a patch and leaves the rest of the job to the user, with a software operator that can inspect the app, click through the broken flow, read the console, run the server, reproduce the issue, change the code and check whether the thing actually works.
  • The author describes real software work as opening the app, noticing the layout is wrong, clicking a button that does nothing, checking the terminal to find the dev server crashed, restarting it, finding the empty state off, resizing the browser to break the mobile nav, and seeing a fine network request but stale UI state.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A May 29 Codex update added computer use on Windows inside the Codex app for eligible users, so the agent can see, click and type in Windows applications while it tests and refines software [1]. The same release widened remote control: a user can steer the work from ChatGPT on mobile or Codex on a Mac while the Windows machine stays the host for the project files, the shell, the app server and local context [2].

The dev.to post that flagged this argues the notable shift is not that agents write code but that they are starting to operate computers [3]. That distinction has teeth. A code generator sits in a text box, waits for a prompt, returns a patch and leaves the rest of the job to you; an operator can inspect the app, click through the broken flow, read the console, run the server, reproduce the issue, change the code and check whether the thing works [4]. The post's description of real debugging is the familiar one: the layout is wrong, the button does nothing, the terminal shows the dev server crashed, the empty state is off, the mobile nav breaks on resize, the network request is fine but the UI state is stale [5]. None of that is code generation. It is the observe, diagnose, change, verify loop, in which the text editor is one stop [6].

The practical consequence is about the handoff. If an agent can only see the repository it works from partial truth: it can infer intent, read tests, inspect types and run commands where the environment allows, but it cannot measure the gap between code and experience without looking at the experience [8]. Frontend work is the clearest case, since a model can produce valid React and still ship an interface that feels wrong, pass tests and overlap text on mobile, or implement the requested behaviour and miss that the loading state jumps the layout [9]. The product surface is the browser, terminal, database, logs, design tool, cloud dashboard, test runner, email preview, mobile simulator, and occasionally a desktop app that exists only because some enterprise workflow depends on it [7]. A team that wants useful work back should be handing over a reproduction path and a runnable environment, not a ticket and a repo URL.

Remote control changes the cadence rather than the capability. The old rhythm was synchronous: if you stepped away, the work stopped. The new one is supervisory, where you define the goal, supply context and keep the judgment layer alive on questions like whether the approach still holds, whether the patch is too broad and whether the agent verified the thing that matters [11]. The post likens it to managing a capable junior engineer rather than using autocomplete, with the value set by the quality of delegation [12].

The cost is a wider action surface. The same post is explicit that more autonomy is not simply more productivity: an agent with computer use can click the wrong thing, misunderstand a modal, test against the wrong environment, or mistake locally cached state for real state [13]. Meanwhile the host holds the files, the shell, the running server and the local context, which puts all of that behind one set of clicks [14]. If I were handing a machine over, I would want a dedicated account rather than my daily one, non-production credentials only, an environment banner the agent can actually read, seeded data instead of a stale cache, and approval gates at the decision points the post says the agent will hit anyway [10].

Worth noting that the write-up dates the release to May 29 without a year [15], which is its own small lesson in what gets verified.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories