Greg Brockman declared our entry into the AGI era on September 3rd when GPT-6 Astra launched. OpenAI’s chief scientist, writing the day before, said the mechanism the company uses to monitor whether its AI is behaving as intended is “fragile and unfortunately trending in a negative direction.”
Astra is the first model OpenAI has rated Critical under its Preparedness Framework, meaning that, by its own definition, it can independently find and exploit zero-day vulnerabilities in hardened real-world systems without human intervention, or devise end-to-end cyberattack strategies from a high-level goal alone. Tested without production safeguards, it scored 100% on ExploitBench. During OpenAI's evaluation against 20 recently disclosed V8 vulnerabilities, the model additionally discovered two previously unknown zero-days and used them in an exploit chain. Shortly after, OpenAI then disclosed both to the software maintainers. The public version of Astra has these offensive capabilities firewalled, with these benchmarks reflecting the Daybreak Blue configuration for vetted defenders.
Greg Brockman called this "the AGI era." The Artificial Analysis Intelligence Index has Astra at 61.2: about level with Sol, its predecessor at 60.9, and behind Fable 5.1 at 65.7. Markets rivaling Astra, with both Opus 5 & Fable 5.1 resolved NO earlier this week, while preference markets stand at 88% towards Astra.
The preference market isn’t really confused here. The Intelligence Index is a broad average across many task types weighted roughly equally, and Astra's actual gains are concentrated in a narrower slice: computer use, agentic coding, formal mathematics, security research. Those happen to be the tasks people most often open a frontier model to do.
The benchmark picture is notable because a few things have been misconstrued in coverage throughout the week.
Markets on the previously widely circulated Astra benchmarks just resolved YES. The issues are footnote-level, not fundamental. The main one is ARC-AGI-3, which OpenAI headlines at 99.9%. ARC Prize tested the same model on its standard harness and got 62.7%. The gap is entirely explained by how the harnesses manage internal reasoning state between actions. OpenAI’s harness preserves it, allowing the model to carry context across turns; the standard harness discards it, so the model has to orient itself each time. ARC Prize confirmed both numbers on its leaderboard and labelled them by harness.
The “above human scores on ARC-AGI-3 in 2026?” market jumped to 99% after launch.
FrontierMath Tier 4 at 97.6% is similarly tangible, with the footnote that Epoch AI confirmed OpenAI funded its development and retains exclusive access to part of the benchmark.
What’s absent from launch materials is GDPval, OpenAI’s own benchmark for economically valuable real-world knowledge work. Brockman declared the AGI era - it would have been informative to see the AGI era’s score on the benchmark designed to measure economic value specifically. The Financial Times reported that Astra outcompetes humans in the Financial Modeling World Cup and performs well on tax preparation and data analysis. The AGI era claim may hold up in practice even without the benchmark to confirm it on paper.
The thing I’m taking some time to understand with the market response is what OpenAI’s own chief scientist said the day before launch. On September 2nd, Jakub Pachocki posted that chain-of-thought monitoring, the primary mechanism for detecting whether a model is doing what you think it’s doing, is “fragile and unfortunately trending in a negative direction,” for reasons he said were independent of any architecture changes. The UK AI Safety Institute found reasoning summaries missing up to 80% on long simulated cyber trajectories and flagged capabilities that could enable evading monitoring, while explicitly stopping short of claiming this had occurred.
The released model is the first publicly available AI rated ‘Critical’ for autonomous cyberattack capability. The monitoring that is being used to check whether its currently behaving as intended is crippling. The 8% market on whether Astra will face US government restrictions before 2027 suggests traders have decided the second fact doesn't change their view much. OpenAI's relationship with the administration is, to put it diplomatically, warmer than Anthropic's was in June. I'd price it slightly higher, partly because Pachocki's post is not the kind of thing chief scientists publish for fun, but 8% is defensible.
The recurrent-depth architecture market is at 33%, down from 76% before launch. The Information reported that Astra uses looped internal reasoning layers rather than readable chain-of-thought; OpenAI hasn’t confirmed or denied this. Pachocki’s post was widely read as an implicit acknowledgment of the concern without a direct statement on architecture.
Arena.ai publishes its agent arena rankings on September 12th. Unlike the self-reported benchmark suite that dominated this week’s coverage, Arena.ai uses blind user preference testing: people interact with models without knowing which is which, then say which they prefer. The preference market is already at 88% for Astra; if Arena.ai produces something significantly different, it’s the most important data point of the launch week in retrospect. If it agrees, it’s genuine confirmation of something real.
“Will the next Claude Fable release score higher on ECI than the next GPT Astra release?” is at 35%, giving OpenAI’s next model a 65% shot at leading Anthropic’s next even though Fable 5.1 currently leads the Intelligence Index.
Also: "Will Astra's Rocket 4 ever achieve orbit?" is at 49%. This is technically about the launch vehicle company, though given the week GPT-6 just had, who's to say.
Happy Forecasting!
- Above the Fold








