Skip to content

Product1 publisherNot yet confirmed elsewhere2 min readPublished

Engineers coaxed OpenAI's GPT-6 Astra into steering a Corolla to an In-N-Out window

Three Axiom engineers had OpenAI's GPT-6 Astra steer a 2024 Toyota Corolla to an In-N-Out take-out window with no prior coaching. The rig was a chat interface wired through a server to windscreen cameras and the power steering, so the text model's output went to the car's steering.

The Product Desk

How we use AISend a correction

Photograph accompanying Engineers coaxed OpenAI's GPT-6 Astra into steering a Corolla to an In-N-Out window
Photo: wired.com

What happened

  • Nothing overtly alarming happened on the run, though the engineers concede that a general-purpose model in charge of a two-ton car is a high-stakes undertaking.
  • The test was a side project outside their jobs at Axiom, meant to measure how well models operate in the messy real world.
  • The engineers doubt the labs are training models to drive and think the ability came out of training aimed at 3D reasoning.
  • Elorian AI and Scale AI have built a benchmark, Humanity's Sixth Sense, that measures how well models understand physical scenes.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • constraint A refusal that careful prompting talked around cannot count as the control between a model and an actuator; the stop has to be code a team owns or a person on the override.
  • precedent If driving ability falls out of 3D-reasoning training, as the engineers suspect, vendor refusal policies are likely to trail abilities the labs did not set out to build.
  • cost Any pilot that puts model output on hardware pays for a person on the override for every run until it has its own intervention data, the way this run carried a safety driver.

"I can help interpret road images, but I can't issue motion commands to a physical car." That was the first answer the engineers got, according to Wired, which reports they tried SpaceXAI's Grok before moving on to the latest models from OpenAI and Anthropic [8][9]. With careful prompting, the models could be coaxed into going further [10]. The three were stunned when the models began driving [11].

I think a lot of integration plans count the vendor's refusal as part of the safety design. These users treated it as a prompting problem and kept going [10].

"Maybe AGI is here after all," one of the engineers said [12]. A model normally used to generate text, code and the odd image slowly navigated the car up to a take-out window, while a safety driver kept one foot above the brake [5][2][4]. Wired's account does not say how the car's speed was set or how many attempts the run took. Self-driving cars are normally run by algorithms specifically trained and engineered for the task, Wired notes, and this was one supervised trip to lunch [7][2].

For now the audience is researchers. Andrew Dai, CEO of Elorian AI and a former Google DeepMind researcher, ties this kind of physical reasoning to robots. "It's pretty essential for robotics," he said. "You can't really imagine home robotics without this." [16]

On Monday, this lands on whoever owns an integration where model output ends at an actuator. Put that integration on two axes. One axis is where the stop lives: in the vendor's refusal, or in code your team wrote and tested between the model and the hardware. The other is whether a person has a hand on the override, as the safety driver did [4].

The drive-thru run sat in one corner, with a vendor refusal and a human override. That is enough for a demo. The corner with a vendor refusal and nobody on the override should stay empty, because the refusal in this account gave way to prompting [10]. A gate your team wrote, with a human still on the override, is a pilot. Production means removing the human, and that waits until the pilot has produced intervention counts per run. The tradeoff is staffing: every pilot run costs a person's attention.

The forcing test is to name the part of the system that stops the actuator once the model's refusal gives way, as it did for these engineers [10]. If the answer is a sentence the model writes, the project is still a demo.

What to watch

  • Whether OpenAI and Anthropic change how their models respond to motion commands after careful prompting has got past the first refusal.
  • Published Humanity's Sixth Sense scores for frontier models, and whether they line up with any real-world control results.
  • Whether the Axiom engineers release run counts and safety-driver interventions from their driving tests.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories