OpenAI dropped a big one this week: GPT-6 Astra, and they're not being shy about it. Greg Brockman, OpenAI's president, literally opened the announcement with "welcome to the AGI era" — and said he personally believes they've crossed that line, even if he's leaving it to the rest of us to decide whether "Astra" fits the definition.
That's a bold thing to say out loud. So I dug into what's actually new here, past the headline.
What Astra is actually better at
This isn't a "slightly smarter chatbot" update. A few numbers stood out to me:
- Computer use and automation — Astra handles the boring stuff (form filling, CRM updates, calendar management) about 47% faster than the previous model. On the OSWorld benchmark it scores 72.6%, and does it in roughly 40 minutes instead of 75.
- Professional work — it's apparently good at producing documents, slide decks, and spreadsheets that actually follow a template and match a real business standard, and when instructions are vague, it asks a focused question instead of just guessing. That's the part I care about most as someone who edits and designs for a living — "guessing" is usually where AI tools waste your time.
- Coding — 57.9% on Terminal-Bench 4.0, with less back-and-forth needed to get something production-ready.
- Math and science — a 98% score on FrontierMath Tier 4, and OpenAI says it contributed to advancing bounds on prime number gaps — a genuinely old, hard math problem.
- Cybersecurity — 100% on ExploitBench, and during testing it found two previously unknown zero-day vulnerabilities on its own. This is actually the part that made OpenAI nervous enough to label Astra "critical" under their own safety framework — it's the first model to get that label.
The alignment number that's easy to miss
Buried in the announcement is a stat that matters more than it sounds: Astra refuses to carry out unauthorized tasks in 0% of test cases where it should refuse — down from 48% for the previous model. In plain terms, it's much better at not doing things it shouldn't, even when a user pushes it. Amelia Glaese, OpenAI's research VP, put it simply: "When models can do more things autonomously, we have to be able to trust them more."
But it's also harder to watch
Here's the part that isn't just marketing spin: OpenAI's own chief scientist, Jakub Pachocki, acknowledged that Astra is harder to monitor for signs it's working around its own oversight — internally flagged as a "serious" concern. So yes, it's more capable, but the people who built it are also saying, on the record, that keeping an eye on it is getting harder. That's worth sitting with for a second before we get too excited.
Rollout and pricing
Astra is rolling out today to select organizations first, then expanding to ChatGPT Plus, Pro, Business, and Enterprise users, plus the OpenAI API and AWS. If you're building on the API, pricing is $10 per million input tokens and $50 per million output tokens.
My take
This launch lands right in the middle of a genuinely wild week — Anthropic, Meta, and Google all shipped major updates around the same time, so this isn't happening in a vacuum. The competition is clearly accelerating faster than most of us can keep up with, and that's exactly why it matters to actually read what's inside these announcements instead of just reacting to the headline "OpenAI says AGI."
For anyone working with AI tools day to day — writing prompts, building workflows, using AI in your editing or design pipeline — the practical takeaway is this: the gap between "AI that answers questions" and "AI that just does the task for you" keeps shrinking. That changes how we should be prompting, and how much we can realistically hand off. I'll be testing Astra properly once it's available and sharing what I find.
Sources: OpenAI's official announcement, and reporting from CNBC, Axios, and Al Jazeera.


