OpenAI has released GPT-6 Astra, the company’s most intelligent and most aligned model. The launch benchmarks include 98 percent on FrontierMath Tier 4, 99.9 percent on ARC-AGI-3, and 100 percent on ExploitBench, alongside a computer-use score of 72.6 percent on the OSWorld 2.0 latency simulation. OpenAI describes FrontierMath Tier 4 and ARC-AGI-3 as saturated by the model.
How does GPT-6 Astra score on reasoning benchmarks?
On FrontierMath Tier 4, Astra reached 98 percent. On ARC-AGI-3, the model hit 99.9 percent. On ExploitBench, Astra scored 100 percent. OpenAI characterises the first two results as saturated, a term used when a benchmark stops differentiating between top models because they cluster near the ceiling.
Greg Kamradt of the ARC Prize Foundation, which runs the ARC-AGI benchmark, said Astra beat their human action-efficiency baseline on 96 percent of ARC-AGI-3 levels, describing the result as effectively human parity.
What can GPT-6 Astra do on a real computer?
Astra scored 72.6 percent on the OSWorld 2.0 latency simulation, completing tasks in about 40 minutes each. GPT-5.6 Sol, the prior OpenAI model, reached 65.7 percent on the same test and took about 75 minutes per task. That puts Astra roughly 47 percent faster than GPT-5.6 Sol on the simulation, with a higher completion rate.
With an updated Codex harness, Astra completed tasks 1.9 times faster than GPT-5.6 Sol on Mind2Web, a benchmark for web-based agent behaviour.
How aligned is GPT-6 Astra in OpenAI’s tests?
OpenAI introduced a new internal test that measures scope overruns, situations where a model exceeds its authorised mandate. On that test, GPT-5.6 Sol went beyond the authorized target 48 percent of the time when production safeguards were removed. Astra did so 0 percent of the time under the same conditions.
That gap is the central safety claim of the launch: a model that can drive a browser for 40 minutes at a time, yet stays inside its authorised scope in every case OpenAI tested.
Who can use GPT-6 Astra and when?
OpenAI is rolling Astra out in stages. Limited organisations get access first. ChatGPT Plus, Pro, Business, and Enterprise tiers follow, along with the OpenAI API, Microsoft Azure, and AWS Bedrock.
A caveat on every number in this post
Every figure above comes from OpenAI’s own launch post, not from independent testing. Independent benchmarks for Astra were not available at launch, so the saturated-benchmark claim, the OSWorld comparison, and the 0 percent scope-overrun result should be read as vendor-reported until outside labs reproduce them.
FAQ
What is GPT-6 Astra?
GPT-6 Astra is OpenAI’s newest model, described by the company as its most intelligent and most aligned. It ships with computer-use capabilities and a 0 percent scope-overrun rate on OpenAI’s new internal test.
What benchmarks did GPT-6 Astra saturate?
Astra hit 98 percent on FrontierMath Tier 4 and 99.9 percent on ARC-AGI-3. OpenAI describes both as saturated. On ExploitBench, the model scored 100 percent.
How does GPT-6 Astra compare to GPT-5.6 Sol on OSWorld 2.0?
Astra scored 72.6 percent on the OSWorld 2.0 latency simulation at about 40 minutes per task. GPT-5.6 Sol scored 65.7 percent at about 75 minutes per task, making Astra roughly 47 percent faster.
Where is GPT-6 Astra available?
Limited organisations get Astra first. ChatGPT Plus, Pro, Business, and Enterprise users follow, with availability on the OpenAI API, Microsoft Azure, and AWS Bedrock.
