09:00
16:00
09:00
16:00
12:40
11:50
09:00
16:00
09:00
16:00
12:40
11:50
09:00
16:00
09:00
16:00
12:40
11:50
09:00
16:00
09:00
16:00
12:40
11:50
OpenAI has released GPT-6 Astra to paid ChatGPT users, and the first independent benchmarks suggest a more modest upgrade than the company promised. In overall intelligence, the new flagship scored exactly the same as its predecessor.
Sam Altman announced the rollout, which began with Pro, Enterprise and Business Premium before expanding to Plus and Business, leaving Go as the only paid ChatGPT plan without access to Astra.
The model is available in Codex and ChatGPT Work, with Plus and Business users getting roughly 5 to 45 requests every five hours. Pro limits range from 25 to 225 requests, or up to 900 for users on the version with 20x limits. Astra is also available through the API at $10 per million input tokens, $50 per million output tokens and $1 per million cached tokens.
Artificial Analysis, which benchmarks models independently using a standardized public methodology, gave Astra 61 points on its Intelligence Index — exactly the same score as GPT-5.6 Sol. The result also leaves OpenAI’s new flagship behind Anthropic’s Claude Fable 5.1 at 66 points and Meta’s Muse Spark 1.3 at 62.

Astra performed better in coding, scoring 67 on the Coding Agent Index when running in Codex and landing roughly alongside Claude Opus 5 and Fable 5. Fable 5.1 in Claude Code still holds the top spot with 70 points.

The new model is 2.5 times more expensive per token than GPT-5.6 Sol, although it uses roughly a third as many tokens. That efficiency largely offsets the higher price for coding at maximum reasoning effort, bringing the cost per task close to its predecessor. The difference is more pronounced on general intelligence tasks, where Artificial Analysis found Astra to be around 75% more expensive per task than Claude Fable 5.
Results across individual benchmarks were similarly uneven. Astra cut its hallucination rate on AA-Omniscience from 92% to 51%, improved on large projects involving thousands of files and made smaller gains in a benchmark focused on mathematics and coding. At the same time, it scored lower on practical tasks covering 44 professions and dropped by two to three points in several tests, including customer support and large-document analysis.
Overall, the independent results show the clearest gains in coding, where Astra combines stronger performance with much lower token usage. In general intelligence, however, it has yet to overtake Anthropic’s leading models, and the numbers so far fall short of the revolutionary leap OpenAI promised.

