OpenAI launches GPT-6 Astra, a premium reasoning model aimed at the top of the leaderboard

High score, higher price
GPT-6 Astra (max) arrived on September 3, 2026 with a score of 61 on the Artificial Analysis intelligence index, nearly double the median of 36 for models in its price tier. The index blends reasoning, knowledge, math and coding benchmarks, and OpenAI's new model finished the evaluation with 42 million output tokens — below the 62 million median — suggesting relatively concise answers. Running the full benchmark cost $3,013, a figure that underscores the expense of testing a model at this level.
Pricing that breaks the scale
Astra sits at the top end of the pricing spectrum: $10 per million input tokens and $50 per million output tokens, against medians of $2 and $10 respectively. Under a weighted pricing model (7:2:1 hit-cache/input/output) the blended rate falls to $7.70 per million tokens, still well above competitors in the same band, which sit above $1 per million. OpenAI has not disclosed the parameter count, and the weights remain closed; this is a fully proprietary model, neither "open weights" nor open source.
Multimodality and a wider context window
The model accepts text and image input, produces text only, and ships with a 1 million token context window — a substantial jump from the previous generation. Its training knowledge cuts off in April 2026. As a reasoning model, Astra runs an extended chain-of-thought before answering, which partly explains the elevated token volume and the corresponding cost.
What remains unknown, and what it implies
OpenAI has not revealed the model's size in parameters, has not published results on other standard benchmarks (MMLU, GPQA, HumanEval), and has not provided a direct comparison with Claude 4 Opus or Gemini 2.5 Pro. The Artificial Analysis index is currently the sole point of comparison, and the evaluation was conducted under lab conditions, not in production. Without supplementary data, it is difficult to judge whether the price premium is justified for specific tasks or mainly reflects branding and positioning in the premium segment.