Its models are off-the-shelf frontier systems, used unmodified, and the model itself is the driver.
That distinction matters because most earlier evaluations of language models on driving have been exercises in question answering or simulation.
Each model ran inside its own native agent environment at medium reasoning effort: Astra and GPT-5.6 Sol in Codex, Claude Fable 5.1 in Claude Code, and Grok 4.6 in Cursor.
The researchers noted that models served with lower latency and higher throughput could have an edge for exactly that reason.
After its first run, Astra reflected that it had declared the car aligned too early, and it resolved to use shorter movements at lower speeds near bends and islands.