Build1 distinct publisher3 min readPublished
Two additions push a coding agent past editing source, into watching a running application and auditing a repository. Both rely on access that two documented flaws have already abused.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The pane opens on the agent's action. Claude builds, launches or checks an application in Apple's simulator, and the device screen starts streaming beside the conversation [2]. That covers a failure the compiler stays silent about, where interface code builds cleanly and the layout still breaks or a control does nothing [3]. Anthropic's framing is that the agent can compare its implementation against the running application before presenting the work for review [4]. Everything the check needs sits on one machine, because the pane depends on a local Xcode install with the iOS platform present [6].
The security plugin is the half I would test before trusting. Anthropic's release documentation describes a set of agents that map the repository, produce a threat model, hunt for vulnerabilities and then review each other's findings [8]. That is a workflow description with no detection rate and no false-positive rate attached. For the pipeline to be worth its triage cost in your repository, two things have to hold: the architecture map has to resemble the system you think you have, and somebody has to read every finding it emits. Neither is in the notes. The immediately useful knob is scope, since a scan can be pointed at a branch diff, a pull request or a single commit instead of the whole tree [9]. Per-commit is the size that fits inside code review.
Restricting three commands to explicit user invocation is a capability being taken back. Those rarely get shipped [11]. Set beside the review gate on patches, the pattern is consistent: the agent produces candidates, a person promotes them.
The cost of all of this is access. A coding agent that is useful needs permission to read repositories, execute commands and interact with development environments [23]. In a GitHub security advisory published on September 9, 2025, Anthropic described a high-severity flaw in Claude Code releases earlier than 1.0.105, in which a malicious Git user.email value could trigger arbitrary code execution before workspace trust was accepted [15]. On February 25, 2026, Check Point Research said malicious Claude Code project configurations could enable remote code execution and theft of Anthropic API credentials, by way of hooks, MCP servers and repository-controlled environment variables [16]. Those two dates are 169 days apart [20]. Both start in the same place: content the repository supplies being read as configuration. A commit-author field that executes code was on nobody's roadmap.
For scale, OpenAI has described its Codex desktop app as a command center that supports delegating long-running work to multiple agents, and GitHub has said its Copilot coding agent can work asynchronously in the background and return a draft pull request for review [17][18]. Against those, Anthropic's addition is narrow by its own description, a live view into one Apple development tool on supported Macs [19]. In my context that is the better trade. A visual check on one platform that a human starts is a smaller promise than autonomous background work, and it is a promise you can verify by looking at the screen.
Ranked by verification strength, evidence, and original report placement.
Anthropic's July 20-24 release added two features to Claude Code Desktop: one lets the coding agent inspect a running iOS application, while the other brings vulnerability scans and proposed patches into the coding session.
Anthropic's release notes say that when Claude builds, launches or checks an application in Apple's simulator, the pane opens beside the conversation and streams the device screen live.
A model can generate interface code that compiles while missing layout failures, broken navigation or controls that do nothing.
Simulator access lets Claude compare its implementation with the application's behavior before presenting the work for review.
Anthropic's release notes describe the iOS Simulator pane as a public beta in Claude Code Desktop for macOS on Pro, Max and Team plans.
The iOS Simulator pane requires Xcode with the iOS platform installed and Claude Desktop version 1.24012.0 or later.
Distinct publishers with included, body-backed reporting in this cluster.
runtimewire.com
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
The $559M-versus-$12.3B quarter matters more than the $65B run rate4 distinct publishers
security
Seven AI coding agents run attacker code named in a repository's own .git config2 distinct publishers
build
Claude Code now outruns Copilot roughly two to one in JetBrains' survey of 15,000 developers1 distinct publisher
build
Anthropic's federal ban falls on a record that failed the supply-chain statute1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Precise citations, single relay
Every capability statement here traces to Anthropic's own release notes as summarized by Runtimewire, down to build 1.24012.0 and the four-stage scan, and the two security items trace to an advisory Anthropic filed and a report Check Point published. Those are good documents, cited with dates and version numbers. None of them is checked by a second party anywhere in this coverage, and the claim that most needs checking, that the plugin's findings and patches are any good, is described in architecture and never measured.
Shipped, nobody counted
The features exist and their edges are known: a public beta behind three paid tiers on one operating system, a hard minimum build, and a security plugin you have to install yourself. Past that, silence. No install counts, no named teams, no scan volumes, no benchmark. Public beta plus opt-in installation is precisely the configuration in which real usage is usually smallest, and this reporting offers nothing to argue otherwise.
Restrained on capability, loose on meaning
Give the reporting its due: it states outright that Anthropic published no measurements for scan quality, and it notes that the July release actually took autonomy away, since Claude can no longer call /verify, /code-review or /deep-research for itself. The overreach is in the connective tissue rather than the specifics, a founding safety principle described as translated into product controls, and a scanner whose worth is implied by having four agents rather than by any published result. A live view of one simulator on one operating system is a smaller thing than the framing around it.
Vendor documents, vendor frame
The company shipping the product is the origin of nearly every capability fact, and the founder motivation is sourced to a 2023 Stripe interview rather than to anyone asking a hard question. The single outside voice, Check Point Research, publishes vulnerability findings as advertising for its own security practice and rounds off with a line about collaborating with Anthropic on remediation. That does not make any of it false; it means each fact arrived through a party with a reason to put it out.
Checkable spine, open question
Version strings, plan tiers and disclosure dates are the kind of detail that is embarrassing to get wrong, which raises trust in the factual skeleton. Against that: one publisher relaying documents, a release month whose year we infer from a September publication date, and the question that actually matters to anyone running this, whether the July hardening closes the configuration paths Check Point walked through, left entirely open.