Abstract illustration of a neural network model release

OpenAI ships GPT-6 Astra and calls it the start of the AGI era

OpenAI released GPT-6 Astra this week in a limited preview on September 3, followed a day later by a restricted rollout to paying ChatGPT and Codex users. The company is calling it the largest training run it has ever done, built on more than 100,000 GPUs at its Stargate site in Texas, and is pitching Astra as state of the art across computer use, browsing, software engineering, cybersecurity, and general professional work.

"Welcome to the AGI era"

The line that made headlines came from OpenAI president Greg Brockman, who closed the press briefing with exactly that phrase. Brockman said he personally believes the company has reached AGI, while leaving the label itself up to users to decide. Nvidia CEO Jensen Huang went further, declaring flatly on X that "AGI has arrived" once Astra shipped. Sam Altman, notably, has pushed back on the term as too vague to be useful — a reminder that even inside the companies building these systems, there's no agreed definition of the finish line.

The benchmark dispute

OpenAI's own numbers show Astra scoring 99.9% on a custom run of the ARC-AGI-3 benchmark, a suite specifically designed to resist memorization and brute-force pattern matching. The ARC Prize organization, which maintains the benchmark, disputes that number — its own testing under standard conditions puts Astra at 62.7%. That's still a strong result, but it's a very different story than "basically solved," and it's a useful case study in why benchmark claims from the lab that built the model deserve independent verification before they get repeated as fact.

Why it matters beyond the marketing

Separately, OpenAI has claimed an early Astra-class system cracked a Millennium Prize math problem in 88 hours — a claim that, if it holds up to scrutiny, would be a genuinely significant result independent of the AGI framing. For teams building on top of these models, the practical takeaway isn't the branding fight. It's that the ceiling for agentic coding, long-horizon computer use, and research assistance keeps moving, and eval suites that were sufficient six months ago are already due for a refresh.

← All posts