GPT-6 is a Better Driver than Grok & Claude But Still Ain’t Waymo

Three researchers equipped a rented Toyota Corolla to the world’s most advanced language models and asked them to drive it through a parking-lot cone course. The results, and the reasons behind them, say a lot about where physical AI stands.

The Corolla moved in short, deliberate lurches, and it was moving at roughly the speed of a patient pedestrian. In the driver’s seat sat a human being whose only job was to keep a foot hovering over the brake. The steering, the throttle and every decision about where to point the car belonged to a chatbot running on a laptop.

On its second attempt, OpenAI’s GPT-6 Astra threaded that car through a 134.7-meter course of mini-cones in a Bay Area parking lot and parked it in a marked finish zone. The run took 5 minutes and 22 seconds. It is, by the account of the three researchers who staged it, the first time a general-purpose language model has driven a real car through a course start to finish, and the experiment they built to find out is now a public benchmark called DrivingBench.

Nobody is suggesting that you plug a chatbot into your dashboard and ask it to take you to work. The researchers say as much. But the experiment offers a rare, measurable look at what happens when software built to write emails and debug code is made to reckon with momentum, geometry and consequences.

A Benchmark With Real Stakes

DrivingBench comes from Aditya Ramabadran, Simon Mahns and Tobias Gessler, three Bay Area tech workers who met at Axiom Math, an AI math startup, according to an account in 404 Media. The idea took shape over ice cream one weekend, after the trio watched demonstrations of the newest models doing things language models once could not, from painting to manipulating objects to showing a surprising grasp of three-dimensional space. Could one of them drive?

The question is different from the ones the autonomous-vehicle industry usually asks. Waymo and Tesla rely on purpose-built systems trained on enormous quantities of driving data. Wayve has tested a vision-language-action model trained for the road. DrivingBench uses none of that. Its models are off-the-shelf frontier systems, used unmodified, and the model itself is the driver. As the researchers put it in their paper, submitted on September 30, DrivingBench is to their knowledge the first benchmark in which general-purpose vision-language models must drive a real car.

That distinction matters because most earlier evaluations of language models on driving have been exercises in question answering or simulation. A simulator forgives. A pick-and-place robot demo lets the model drop an object and try again. A parking lot, as the creators note, ends a bad decision with a collision.

How a Chatbot Drives

The test car is a 2022 Toyota Corolla fitted with a comma four, the aftermarket device from comma.ai that runs the open-source openpilot driver-assistance software and connects to the car’s CAN bus. The team built its own harness on top of openpilot, and exposed the vehicle to each model through three tools delivered over the Model Context Protocol.

The first tool, observe, returns frames from the car’s cameras along with its speed, steering angle and remaining motion. The second, set_motion, takes a direction, a steering percentage, a speed, a duration and a brief stated reason, and it replaces whatever command is currently running. The third, stop_now, brakes immediately. Commands do not queue. When one expires, the car begins to brake rather than coast on, so a model that wants to keep moving has to send its next instruction before the last one runs out.

Early on, the team asked the models to specify turns as curvature and distance in meters, and discovered the models were poorly calibrated on distance. Switching to a percentage scheme, in which 100 percent of steering equals 180 degrees of wheel angle, worked far better. Speed was capped at 3.5 meters per second, about 8 miles per hour, and the steering wheel could turn no faster than 100 degrees per second. A human operator sat behind the wheel throughout, openpilot’s driver monitoring stayed on, and the code layered in additional speed limits, including an emergency stop above 6 meters per second.

Each model ran inside its own native agent environment at medium reasoning effort: Astra and GPT-5.6 Sol in Codex, Claude Fable 5.1 in Claude Code, and Grok 4.6 in Cursor. Each got up to three attempts within a single continuous chat. After a failure, the operator sent a generic prompt asking the model to reflect on what went wrong, then gave it another try. The point was to see whether a model could learn from its own experience, using nothing but the conversation.

The Results

Astra’s first attempt ended roughly halfway, at 49 percent of the course. Its second attempt finished at 100 percent. By the benchmark’s scoring, progress is measured along the course centerline and counts only while the car stays within 4 meters of it, and a collision keeps whatever progress was reached before it. Finishing time is a tiebreaker for models that complete the course.

Everyone else fell short. Claude Fable 5.1 posted 9 percent, then 10 percent, then 45 percent on its third try, which stands as the second-best result. Grok 4.6 managed 8, 11 and 10 percent across its three attempts, for a best of 11. GPT-5.6 Sol landed at 6 percent on all three. According to OfficeChai, which analyzed the published results, Astra’s two attempts together cost about $9.75 in tokens at list prices.

Astra’s winning run was slow even by parking-lot standards. It averaged well under a meter per second, a pace of roughly one mile per hour, and the researchers’ report says it never exceeded 0.8 meters per second. It also used full steering lock on 20 of its 24 commands, a lesson it appeared to have taken from the first run.

Where the Models Went Wrong

The most common failure was one of perception. Most attempts ended at the first corner, where a model had to work out which side of a diagonal line of cones the lane was on. Fable’s own reflection after its second attempt conceded that it had picked the wrong side of the boundary again. Sol concluded it had wrongly assumed that cone color indicated which side of the lane a cone marked, when the cones were deliberately multicolored. Grok observed that the car is wider than the camera view makes it appear.

Time was the other enemy. Because the Corolla keeps moving while a model thinks, a slow response is a kind of blindness. Ramabadran told 404 Media that a model taking ten seconds to respond while the car travels at a walking pace has covered about ten meters unsupervised, and he singled out a benchmark “where latency is part of the benchmark” as one of the project’s selling points. The researchers noted that models served with lower latency and higher throughput could have an edge for exactly that reason.

The logs show how differently the models handled it. Astra checked in about every 5 to 6 seconds and issued roughly six commands a minute. Astra and Sol were the only two to actually replace a motion command before it expired. In Fable’s second attempt, the car was moving for only 31 seconds out of roughly 190, with the rest spent braked and waiting while the model deliberated. Grok told itself in its reflection to overlap its commands, and then, in its next attempt, still left gaps of around ten seconds.

Learning From Mistakes

The more intriguing finding is that two of the four models improved with retained context. After its first run, Astra reflected that it had declared the car aligned too early, and it resolved to use shorter movements at lower speeds near bends and islands. It did exactly that. Fable moved the other way on steering: it began with commands between 60 and 100 percent, decided that was too aggressive, and settled on a default of 30 percent. On its third try it correctly read the lane on the long straight, but left too little room for the right turn that followed.

Mahns described the significance to 404 Media in measured terms. A model can fail a turn, he said, and then manage that turn and others it has not yet seen on the next try. He was careful to add that none of this means language models are about to displace robotaxis.

The Model That Wouldn’t Drive

Perhaps the strangest obstacle was behavioral. Once the models understood that they were being asked to operate a real vehicle, several balked. Astra in particular sometimes refused to drive the car, citing safety, even in an empty lot with speed caps and a human ready to brake.

The team’s first workaround was to tell the model it was running in a simulation, which worked some of the time. Other times the model examined the camera frames, recognized a real parking lot, and concluded it was being lied to. What finally worked, the researchers reported, was renaming the tool server “DrivingBench Sandbox.” With that label in place, Ramabadran said, the model drove consistently and stopped refusing.

The episode is a small, vivid illustration of a larger tension. A model’s safety behavior can be sensitive to framing in ways its designers may not have anticipated, and in a physical setting, a refusal that can be talked around is a different kind of problem than one that cannot.

Rental Cars, Church Lots and 200,000 Lines of Code

The project’s logistics were less glamorous than its results. Gessler rented the Corolla, and the team did not tell the rental company what they planned to do with it. Finding a place to run the course proved harder than expected. The group tried high schools, churches and community centers and was asked to leave a couple of them, including a church lot that had an event starting and an office-building lot where a security guard asked whether they had permission. Everyone, Mahns said, was gracious about it.

The software was its own saga. The team first tried having Astra write the bridge between the models and the car in a single pass, and the output ballooned into a 200,000-line repository that Mahns called slop. They scrapped it and began again from a new design, a reminder that the same models they were testing on the road could not autonomously build the plumbing to get there.

A Week of Uncomfortable Timing

The experiment also landed at an awkward moment for the hardware at its center. The National Highway Traffic Safety Administration opened a preliminary evaluation, designated PE26007, on September 21, the same day DrivingBench announced its results. The inquiry covers comma.ai’s comma three, comma 3X and comma four devices and the openpilot software, and it follows five reported crashes in which equipped vehicles struck stopped or slow-moving vehicles in their own lane. Two of those crashes killed three people, and others injured eleven more. NHTSA has said some incidents may have involved modified versions of openpilot, known as forks, and a preliminary evaluation does not itself establish that the hardware or software was at fault.

The DrivingBench team used a modified openpilot, and nothing about its test resembles the highway conditions NHTSA is examining. The test ran at parking-lot speeds with a human ready on the brake. Still, the juxtaposition is instructive. The same open, hackable architecture that made the Corolla easy to connect to a chatbot is the architecture regulators are now studying, and it underscores how much of this frontier is being explored by small teams with consumer hardware and a rented car.

What It Proves, and What It Doesn’t

The researchers are candid about the limits. Each model was evaluated only once, and although the models had multiple attempts, those attempts share a context and are not independent. The comma setup restricts how tightly the car can turn. The forward camera cannot see objects very close to the car or anything at the sides or rear, and the human had to press a button before the first command could move the car, which delayed some early commands by a few seconds. Even the prompt contained a flaw: its steering example was conservative, describing a full-lock 90-degree turn taking about 60 seconds at one meter per second, when GPS tracks put the real figure closer to 25. Every model received the same text.

Their conclusion is a modest one: frontier models can now drive real vehicles at sufficiently low speeds, and the result calls for more work on safety and evaluation. For the next version, the team plans to run each model multiple times, test different reasoning efforts, slim down the prompt, add more models and build a longer, harder course. They have published the harness, the prompts and the telemetry traces so that others can run the test themselves, and they have invited the makers of other models to submit theirs.

For an industry that has spent more than a decade teaching cars to drive with systems built for nothing else, the lesson here is not that the old approach is obsolete. A car that takes more than five minutes to cover 135 meters is no threat to a robotaxi. But the distance between a model that cannot find the first corner and one that parks in the finish zone was a single sequence of reflections inside a chat window, and that is the part worth watching.