GPT-6 Astra Fully Rolled Out to Premium Users, Performance Benchmarks Released
nashnova research
OpenAI opened GPT-6 Astra to all premium subscribers just two days after its staged launch; benchmarks show it leads on desktop-operation tasks but trails Claude Fable 5.1 on composite intelligence scoring, signaling the flagship-model race is far from settled.
Why did OpenAI rush to full availability in two days?
On September 3, Pro users hit queue problems and pushed back hard. CEO Sam Altman publicly apologized on X, calling it "a messy launch."
Two days later, on September 5, Astra went live for all Pro, Enterprise, and Business Premium users — accessible in ChatGPT Work and Codex, with API access simultaneous.
This means → the accelerated rollout was a direct response to backlash, not a planned timeline. Plus and standard Business users still face a wait of several more days.
How much better is Astra than the previous generation?
On OSWorld 2.0 — a benchmark simulating human desktop operations — Astra scored 72.6% versus 65.7% for GPT-5.6 Sol, a gain of roughly 7 percentage points.
Average time per task dropped from 75 minutes to 40 minutes, with accuracy and speed improving in tandem.
In plain terms = at the specific job of controlling a computer like a human, Astra is meaningfully faster and more accurate — but that is only one dimension.
Why does its composite score still trail Claude?
On Artificial Analysis's independent composite intelligence index, Astra scored 61.2. Anthropic's Claude Fable 5.1, released the same week, scored 65.7 — a 4.5-point gap.
The two benchmarks measure different things: OSWorld tests "hands-on operation"; the composite index tests "broad reasoning." Astra leads on the former and trails on the latter.
This reflects a broader reality: the AI flagship race is no longer won by a single leaderboard — every player has strengths and blind spots.
What have developers actually built with Astra?
Developer Matt Shumer used Astra to construct a Manhattan street-block scene in Unreal Engine over one week, using a "manager loop" architecture: one Astra instance broke down the task list while another executed, orchestrating up to 96 sub-agents in parallel.
Developer Tom Krcha fed Astra an old steam-train blueprint; within minutes it generated 3,295 editable objects in Blender. The demo drew over 750,000 views on X.
Immunologist Derya Unutmaz entered a single prompt. Astra autonomously wrote the narration, produced animations, and generated images, outputting a T-cell educational video. The scientist — who has studied T cells for over 35 years — said the result exceeded his expectations.
Why is OpenAI telling developers to subtract, not add?
In a model guide released alongside the full rollout, OpenAI engineer Victor Nunez advised developers to audit every existing AGENTS.md and Skills file, labeling the recommendation "strongly advised."
The reason: Astra is more sensitive to instructions than its predecessors. Vague rules trigger repeated confirmations; conflicting rules cause task interruptions.
In plain terms = the patch-upon-patch prompts developers accumulated to compensate for weaker models may now actively interfere with Astra. More instructions are not better — more precise instructions are.
What comes next?
The 4.5-point gap between Astra's composite intelligence score and Claude Fable 5.1 is the key metric to watch in subsequent releases.
This means → for investors, leading on one benchmark does not equal product leadership; the composite-capability gap is the variable that shapes market positioning.
市场有风险,内容仅供研究参考,不构成投资建议。