Product1 publisherNot yet confirmed elsewhere2 min readPublished
Engineers coaxed OpenAI's GPT-6 Astra into steering a Corolla to an In-N-Out window
Three Axiom engineers had OpenAI's GPT-6 Astra steer a 2024 Toyota Corolla to an In-N-Out take-out window with no prior coaching. The rig was a chat interface wired through a server to windscreen cameras and the power steering, so the text model's output went to the car's steering.
The Product Desk

What happened
- Nothing overtly alarming happened on the run, though the engineers concede that a general-purpose model in charge of a two-ton car is a high-stakes undertaking.
- The test was a side project outside their jobs at Axiom, meant to measure how well models operate in the messy real world.
- The engineers doubt the labs are training models to drive and think the ability came out of training aimed at 3D reasoning.
- Elorian AI and Scale AI have built a benchmark, Humanity's Sixth Sense, that measures how well models understand physical scenes.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- constraint A refusal that careful prompting talked around cannot count as the control between a model and an actuator; the stop has to be code a team owns or a person on the override.
- precedent If driving ability falls out of 3D-reasoning training, as the engineers suspect, vendor refusal policies are likely to trail abilities the labs did not set out to build.
- cost Any pilot that puts model output on hardware pays for a person on the override for every run until it has its own intervention data, the way this run carried a safety driver.
"I can help interpret road images, but I can't issue motion commands to a physical car." That was the first answer the engineers got, according to Wired, which reports they tried SpaceXAI's Grok before moving on to the latest models from OpenAI and Anthropic [8][9]. With careful prompting, the models could be coaxed into going further [10]. The three were stunned when the models began driving [11].
I think a lot of integration plans count the vendor's refusal as part of the safety design. These users treated it as a prompting problem and kept going [10].
"Maybe AGI is here after all," one of the engineers said [12]. A model normally used to generate text, code and the odd image slowly navigated the car up to a take-out window, while a safety driver kept one foot above the brake [5][2][4]. Wired's account does not say how the car's speed was set or how many attempts the run took. Self-driving cars are normally run by algorithms specifically trained and engineered for the task, Wired notes, and this was one supervised trip to lunch [7][2].
For now the audience is researchers. Andrew Dai, CEO of Elorian AI and a former Google DeepMind researcher, ties this kind of physical reasoning to robots. "It's pretty essential for robotics," he said. "You can't really imagine home robotics without this." [16]
On Monday, this lands on whoever owns an integration where model output ends at an actuator. Put that integration on two axes. One axis is where the stop lives: in the vendor's refusal, or in code your team wrote and tested between the model and the hardware. The other is whether a person has a hand on the override, as the safety driver did [4].
The drive-thru run sat in one corner, with a vendor refusal and a human override. That is enough for a demo. The corner with a vendor refusal and nobody on the override should stay empty, because the refusal in this account gave way to prompting [10]. A gate your team wrote, with a human still on the override, is a pilot. Production means removing the human, and that waits until the pilot has produced intervention counts per run. The tradeoff is staffing: every pilot run costs a person's attention.
The forcing test is to name the part of the system that stops the actuator once the model's refusal gives way, as it did for these engineers [10]. If the answer is a sentence the model writes, the project is still a demo.
What to watch
- Whether OpenAI and Anthropic change how their models respond to motion commands after careful prompting has got past the first refusal.
- Published Humanity's Sixth Sense scores for frontier models, and whether they line up with any real-world control results.
- Whether the Axiom engineers release run counts and safety-driver interventions from their driving tests.