GPT-6 Astra shows large gains on new robotics spatial reasoning benchmark
In early testing on a new robotics benchmark called StationeryBench, GPT-6 Astra reportedly completed 7 of 100 tasks using dual-arm robots, while competing model MolmoAct2 finished none. A researcher characterized the results as a marked improvement in spatial reasoning. The figures come from preliminary benchmark runs rather than a full public release.