GPT-6 Astra has gained the ability to drive a car

plurby 285 points 226 comments September 23, 2026
drivingbench.com · View on Hacker News

Discussion Highlights (20 comments)

pietz

Apparently I have a new favorite benchmark. Honestly, this is cool.

mohamedkoubaa

I'd have started with an RC car but to each their own

decodingchris

Super cool benchmark!

amluto

I’m morbidly curious whether the (supposedly) superior compaction support in recent GPT models with an appropriate harness has anything to do with this. A conventional LLM with conventional attention is, of course, wildly unsuitable to continuous tasks like driving, but maybe as the technology advances it will improve in its ability to sort-of work.

Bluestein

Oh, lord. They are going to Jev this.-

worldsavior

5 minutes - 7 dollars.

Mooty

How do they even test this on a model ? I mean it's a multimodal i get that but response time are too big or am i missing something ?

zezcko

I think the most interesting part of this is that Astra initially refused to drive because it realised it was driving a real car and would only obey when the MCP was renamed to DrivingBench Sandbox. This is both an interesting detection by the LLM but also for me an interesting dynamic concerning LLM "jailbreaking". Saying they were driving 7 mph, that it was oversaw by humans and the fact it was an empty course still wasn't enough for the model. The evaluators even tried to convince the model it was a simulation, it STILL wouldn't budge. And yet as soon as the words "bench" and "sandbox" appear, the model apparently sees this as fair game. Is it a known effect that models will be more likely to comply with requests when they're assumed as "benchmarks"?

syntaxing

Surprised they didn’t try Qwen’s recently open sourced driving model https://huggingface.co/Qwen/Qwen-Drive-1.0-4B

valine

The bitter lesson is finally coming for the self-driving cars. The vision stack, 3D maps, lane selection grammar, occupancy networks, it’s maybe all about to give way to a single GPT looking at camera feeds and predicting the next steering wheel adjustment. It’s mostly a latency problem at this point. The models are too big to run locally, but given that open-weight models like Qwen already exist, an open-weight, low latency equivalent to Astra can’t be too far out.

prometheus1992

Wow! but WHY is this a benchmark?? for comparison tesla's model is approximately 10-15B parameter model (estimating from maxxing the hardware that comes with the car at 16gb ram).

blorenz

Pivot this to analyze and coach human drivers to be better drivers.

WarmWash

3.8 flash would be the model to test, it's vision capabilities are excellent (on par with Astra) while also being incredibly fast.

famouswaffles

What did they do to Astra so cracked at vision (and computer use). That ARC 3 score turned out to be no joke/fluke. That huge gap between Astra and Fable (in this case) is basically every hard vison/spatial benchmark i've seen including non-benchmarks like playing games (Portal, Factorio, RimWorld). SpatialBench - https://x.com/spicey_lemonade/status/2096365630190698516 ZeroBench - https://zerobench.github.io/ Robot Arms - https://openai.robocurve.org/gpt-6-astra/

nashashmi

I wonder if the companies would be willing to bet entirely on AI driven innovation if liability for misalignment was put squarely on companies, individuals, compute vendors, and LLM vendors. I don’t think they would opt for it, especially if an alternative option to use human-programmed tech was already available. There is something to be said about emphasizing on liability as a way to freeze or solidify AI Development. Right now it is too unfettered leading to predictions of AI dooms.

soumyadeb

This also explains why Astra is so good at video generation. I have an Astra+Higgsfield setup. I could point it to a Github repo and ask it to generate a product walkthrough and it did a very good job by generating fake screens (e.g. with data filled in) from real ones - which wasn't possible in earlier models

123917

https://x.com/tobiges/status/2098294046469022030 "Sam understands exponentials like no other. During a YC talk last year he predicted that AI would make breakthroughs in science in 2026 and solve a major open problem in 2027. Now here we are..." Now on a new vibe coded website Astra wins the benchmarks ...

mlmonkey

What about Jev? :-D

josefresco

Looks like the "most successful" path drove over empty parking spaces and came close to two curbs?

TomGarden

New pelican on a bicycle? Genuinely though, this is fun but not at all what these models are good for. It's like cooking a meal with your feet or somthing. A youtube challenge video from 2012

Semantic search powered by Rivestack pgvector
7,510 stories · 69,432 chunks indexed