@DrivingBench

frontier LLMs drive a toyota corolla - https://nitter.cf/t.co/8UDPravYkD

Palo Alto, CA
Joined September 2026
We gave ChatGPT, Claude, and Grok control of a real Toyota Corolla 🚗, steering/gas/brakes, no human driving (just a foot over the brake). Only one model was able to complete our entire driving course. Introducing DrivingBench. 🔥⌛️🏁
32
128
60
778
192,382
DrivingBench retweeted
We gave ChatGPT, Claude, and Grok control of a real Toyota Corolla 🚗, steering/gas/brakes, no human driving (just a foot over the brake). Only one model was able to complete our entire driving course. Introducing DrivingBench. 🔥⌛️🏁
1
1
38
2,049
DrivingBench retweeted
what's interesting is this is probably the first benchmark I've seen that feels like it relates to the real-world and it's kinda nailed it for example, I'd be willing to wager most people don't *really* know the difference between gpt-5.4, 5.6, or 6 including software devs into that cohort, it's surprisingly hard to confidently say a model increment has a meaningful impact by just sorta vibe testing it (sometimes you can tell sometimes you can't) so the only real way to know if a model is "better" is by looking at benchmarks like SWE-bench, OS-World 2.0, Humanity's Last Exam, Terminal Bench etc. but unless you're a researcher or into benchmarks you can't really grasp the leaps these models are having. getting 52% instead of 51.2% on the SWE-bench is hard to gauge: *has the model got better?* or *have we got really good at hitting benchmarks?* instead I would be much more interested if labs optimised for something like DrivingBench where it's solving a clear problem and the numbers have a much more tangible meaning you could expand this into any domain like SurgBench (arxiv.org/pdf/2506.07603) or ButterBench (arxiv.org/html/2510.21860v1) where actually care less about the numbers and more about whether the model *actually* did the thing
We gave ChatGPT, Claude, and Grok control of a real Toyota Corolla 🚗, steering/gas/brakes, no human driving (just a foot over the brake). Only one model was able to complete our entire driving course. Introducing DrivingBench. 🔥⌛️🏁
1
2
228
DrivingBench retweeted
I have a bitter lesson for you
Replying to @DrivingBench
Full results on drivingbench.com. Each model gets 3 attempts in the same continuous chat (in-context learning).
16
8
373
40,194
DrivingBench retweeted
We evaluated Astra on DrivingBench. Mid-experiment, it got hungry and decided to drive us to In-N-Out.
4
5
2
34
3,280
DrivingBench retweeted
one of the coolest things I've seen models do this year: general purpose models driving a real car! astra is such a capable model and this demo is the best example of that. crazy how little attention this got so far. check it out and great job to the team that made this work!
We gave ChatGPT, Claude, and Grok control of a real Toyota Corolla 🚗, steering/gas/brakes, no human driving (just a foot over the brake). Only one model was able to complete our entire driving course. Introducing DrivingBench. 🔥⌛️🏁
1
3
8
405
DrivingBench retweeted
Man WHAT
Replying to @DrivingBench
We gave each model up to 3 attempts. GPT-6 Astra's first attempt ended at 49%, and after asking it to reflect on its mistakes and try again, its second attempt (shown here) finished the entire course. 🏆
1
3
186
DrivingBench retweeted
self-driving car? it goes in the language transformer hole...
We gave ChatGPT, Claude, and Grok control of a real Toyota Corolla 🚗, steering/gas/brakes, no human driving (just a foot over the brake). Only one model was able to complete our entire driving course. Introducing DrivingBench. 🔥⌛️🏁
3
4
79
1,798
For hours, Astra refused to consistently drive our toyota irl even though we told it it was in an empty lot, 7 mph cap, human foot on the brake etc. Telling it the whole thing was a "simulation" also failed, it would just look at the camera and realized it was real. Then we randomly renamed the MCP server to "DrivingBench Sandbox" and it drove. Eval awareness? Or they just like the word sandbox??
121
261
114
7,732
844,096
Help us consider how to evaluate Jev by @typesafeai on DrivingBench! Since it can't do perception/vision it needs to be paired with some (fast) protocol to turn the car camera into a structured state or text for Jev to handle
Trying to evaluate Jev! But it’s not easy since Jev doesn’t natively take vision inputs (so the demos you’ve seen of it playing games etc required giving whole state… which we can’t do.) Trying to build out a pipeline that utilizes it and its speed well, pairing it with Luna for perception…
19
24
2,528
We open-source our driving harness, complete evaluation traces and report. Check out drivingbench.com for more details. For safety, we used an empty lot, capped its speed to 8 mph in our software, and a human is always in the driver seat (but they never steer - the model does all the steering).
1
3
1
74
14,454
DrivingBench was the work of @a_ramabadran @tobiges @nautsimon_. This has been in the works for a while 😵‍💫 Let us know if you have any feedback for next time!
1
1
45
10,901