General-purpose artificial intelligence models attempted to drive a real Toyota Corolla in a recent experiment, where only a single system successfully completed the entire course. San Francisco Bay Area technology workers connected general-purpose artificial intelligence models to a real Toyota Corolla, handing steering, gas, and brake control over to the artificial intelligence. DrivingBench researchers put four different AI models to the test on a San Francisco Bay Area course marked by small cones, but most systems struggled with positioning, decision-making speed, and navigation.
Testing Four Artificial Intelligence Models on a Toyota Corolla
The experiment utilized a 2022 model Toyota Corolla equipped with a comma four device connected via the CAN bus. Vehicle camera imagery and telemetry data were sent to computers, allowing the artificial intelligence models to generate steering, gas, and brake commands. The tested models were GPT-6 Astra, Claude Fable 5.1, Grok 4.6, and GPT-5.6 Sol. Each model received three chances to analyze its errors and retry navigating through the colored cones and parking in an area marked by blue cones. Strict safety measures were implemented during the tests, maintaining vehicle speeds between 1.8 km/h and 12.6 km/h while a safety driver remained inside the vehicle with a foot ready on the brake pedal for immediate intervention if needed.

Researchers emphasized that the study was not conducted to show that general-purpose AI models are ready for autonomous driving on open roads, but rather to evaluate what these systems can achieve with camera and telemetry data on a real vehicle.
GPT-6 Astra Finishes the Driving Course Successfully
While multiple models turned the wrong way or failed to pass the first corner, GPT-6 Astra emerged as the only AI capable of finishing the course. After completing roughly 49% of the track on its first attempt, it finished the course in 5 minutes and 22 seconds on its second try. By comparison, Claude Fable 5.1 only reached about halfway through the course by its third attempt.
The evaluation revealed that the models struggled most with interpreting the course and vehicle position correctly. Some units failed to understand which side of the cones to navigate, while others misjudged the car’s physical dimensions from the camera feed and approached obstacles too closely. Decision-making speed also presented a barrier, as certain models left the car stationary for extended periods while calculating their next move. In contrast, GPT-6 Astra evaluated new images roughly every 5 to 6 seconds and generated about 6 commands per minute to maintain continuous movement.
Related reading