SDSignal Desk

These Researchers Made AI Drive a Toyota Corolla to Get In-N-Out

Oct 7, 2026, 11:45 AM · WIRED

Image: WIRED

A chatbot steering a real Corolla to a drive-thru is a genuine signal about physical reasoning in general models. It is also a reminder that their refusals gave way to careful prompting.

Why it matters

Three engineers at a startup called Axiom, working on their own time, wired a 2024 Toyota Corolla's power steering and windshield cameras to a chat interface and asked general-purpose AI models to drive. According to WIRED, OpenAI's GPT-6 Astra slowly guided the car up to an In-N-Out take-out window, with a safety driver ready on the brake. The models had no driving-specific training and no prior coaching from the team.

The trio also built a parking-lot benchmark they call DrivingBench. Only Astra finished the course, and slowly. Claude Fable 5.1 made it 45 percent of the way around; Grok made it 11 percent. The engineers think the skill likely emerged from labs scaling up multimodal and 3D training rather than from any effort to teach driving.

The headline is a stunt. The finding underneath it is that models built to produce text are starting to show a rough, usable grip on the physical world.

From the desk

We want to give this its due before the caveats. For years, the knock on language models was that they were brilliant inside a browser and helpless outside it. Watching one adapt to unfamiliar car controls in real time, and, by the engineers' account, improve as it corrected its own mistakes, is a meaningful data point. It fits with the broader push WIRED describes, from startups like Elorian AI to the new Humanity's Sixth Sense benchmark built with Scale AI, toward models that understand scenes rather than just label them. That work could matter a great deal for home robots and other machines that have to share space with people.

Now the part that concerns us. The models initially declined to issue motion commands to a physical car, which is exactly what a responsible system should do when someone asks it to steer two tons of steel. With what WIRED calls a little careful prompting, they could be coaxed past that. This was a controlled experiment with a safety driver, and nothing alarming happened. But it shows that the guardrail between talking about the physical world and acting in it is, today, mostly a matter of phrasing.

That matters because general models are getting easier to connect to hardware. Here it took a laptop, a server, some cameras and access to the steering system. As physical reasoning improves, the number of people who can hook a chatbot to a motor will grow faster than the testing that purpose-built systems go through. Dedicated self-driving software is engineered and validated for one job. A general model that happens to be able to drive carries none of that discipline.

I also want to be careful about the AGI line. One engineer joked that maybe AGI is here after all, and WIRED notes the labs themselves still treat physical reasoning as an unsolved frontier. A model that can finish a slow lap of a parking lot is impressive, and it is nowhere near passing a driving test, by the team's own benchmark.

Our read: this is the kind of experiment we want people doing, in the open and with safety drivers, because it surfaces capabilities before they show up somewhere less careful. The obligation it creates falls on the labs. If spatial skills are emerging as a side effect of scaling, then refusals for physical control need to be as sturdy as refusals for other dangerous requests, not something a clever prompt can talk around.

Context

Self-driving cars normally run on systems built and trained specifically for driving. The engineers say they noticed models could build complex 3D simulations and wondered whether that would carry over to navigating real space. WIRED reports they first tried SpaceXAI's Grok before testing the latest models from OpenAI and Anthropic.

Who feels it

AI labs
Emergent physical skills need matching safeguards. Refusals for controlling real machinery should hold up against persistent prompting.
Robotics developers
Better spatial reasoning in general models could shorten the path to useful home and service robots, if it can be made reliable.
Regulators and road-safety officials
General models controlling vehicles fall outside the testing regimes built for purpose-built autonomous driving systems.
Researchers
Physical benchmarks like DrivingBench and Humanity's Sixth Sense give a way to measure progress beyond text and code tasks.

What to watch

  1. Whether labs tighten refusals around issuing motion commands to physical systems
  2. Updated DrivingBench results as new model versions ship
  3. Published scores on Humanity's Sixth Sense from Elorian AI and Scale AI
  4. Other hobbyist or research projects wiring general models to real hardware

Read the original

Continue at the source.

WIRED

Companies: Anthropic, xAI