IACSIACSInt'l Academy for Consciousness Studies
AI · News · The Campus Chronicle

OpenAI Declares the 'AGI Era' Arrived With GPT-6 Astra. The Benchmarks Tell a More Complicated Story.

A landmark launch, a triumphant proclamation, and one stubborn leaderboard that Anthropic's Claude Fable 5.1 still tops.

September 16, 2026 · International Academy for Consciousness Studies

OpenAI released GPT-6 Astra on Sept. 3, 2026, calling it the most intelligent and aligned model it has ever built; President Greg Brockman closed a press briefing with the words "Welcome to the AGI era." Brockman took a position that went further than the benchmarks themselves, suggesting that future observers might look back at Astra as the model that marked the arrival of artificial general intelligence, though he acknowledged AGI remains a "gray, fuzzy thing." The model was trained in OpenAI's largest run to date, using more than 100,000 GPUs at its Stargate facility in Texas. The launch came just two days after Anthropic released Claude Fable 5.1 on Sept. 1.

Astra's most striking numbers are its 97.6 percent score on FrontierMath Tier 4, its 99.9 percent on ARC-AGI-3 under OpenAI's own evaluation harness, and a 72.6 percent score on OSWorld 2.0 computer-use tasks, accomplished at roughly 47 percent less time per task than its predecessor, GPT-5.6 Sol. The model is also the first OpenAI system to reach the "Critical" threshold under the company's Preparedness Framework for cybersecurity capability. Yet the scorecard has a notable gap. On Humanity's Last Exam with tools, Astra scores 57.2 percent against Fable 5.1's 65.0 percent, making it the only major academic row Astra loses; crucially, OpenAI's own announcement prose does not mention the result, and on a benchmark named after the end of testing, the new flagship trails every Claude model in the comparison table. On the Artificial Analysis Intelligence Index, an aggregate of general capability, Claude Fable 5.1 posts 65.7, ahead of Astra's 61.2.

OpenAI's headline 99.9 percent ARC-AGI-3 score was itself immediately disputed: the ARC Prize Foundation published a breakdown noting that under standard conditions the score stood at 62.7 percent. Within OpenAI itself there is no unified position; CEO Sam Altman has described AGI as a poorly defined term, while Brockman called the era's arrival his personal opinion. AI evaluation researcher Dr. Rebecca Johnson of the University of Sydney said "OpenAI's own definitions expose serious flaws" in the claim, noting that the company's charter defines AGI as autonomous systems that outperform humans at most economically valuable work, a definition that embeds, in her words, "a value judgement" about what work counts. Researcher Gary Marcus argued that the launch declaration lacked both evidence and a sufficiently precise definition, and said Astra falls short on most of the criteria in his own published framework for general intelligence.

Johnson went further, saying she would "eat my hat if we didn't find trivial things that an eight-year-old can do that Astra fails at." Independent evaluators have also flagged that Astra's 99.9 percent ARC-AGI-3 figure requires a stateful adapter harness; stateless calls produce scores ranging from 17 to 63 percent according to the ARC Prize organization. Humanity's Last Exam, created by the Center for AI Safety and Scale AI, was itself designed as a multi-modal frontier benchmark after standard tests reached saturation, spanning 2,500 questions across mathematics, humanities, and the natural sciences. On that leaderboard, the top three models are now clustered within 0.5 points of each other, suggesting even this harder test is approaching saturation for frontier systems, which raises the deeper question of what meaningful evaluation of general intelligence would even look like next.

When the model that just claimed to open the AGI era cannot top the benchmark designed to be the last benchmark, the most urgent question is not whether AGI is here, but whether anyone in the field can agree on what would prove it.

Sources: 'Welcome to the AGI era': OpenAI launches GPT-6 Astra | VentureBeat · GPT-6 Astra Benchmarks: Is It Really Better Than Fable 5.1? | MindStudio · OpenAI says 'the AGI era' is here. Experts disagree | Information Age | ACS

More in this issue

More from The Campus Chronicle