On September 3 OpenAI released GPT-6 Astra, and Greg Brockman closed the briefing with "Welcome to the AGI era." Three days later Jensen Huang posted on X that "AGI has arrived" — and, in the same tweet, that the next 400,000 GPUs were on the way. A man who sells graphics cards is not what you'd call a disinterested party, but fine.
I read it. Not thrilled, not saddened, mostly I wanted to ask: what, precisely, are we being congratulated on? Because in the thirty-odd years the industry has been saying those three letters, it has never agreed on what they mean.
OpenAI's charter has AGI as "highly autonomous systems that outperform humans at most economically valuable work." So: not "thinks like a human" but "does a human's work." Different things, and the difference between them has suddenly become very expensive. A chess engine has long beaten any grandmaster, a calculator counts faster than Einstein, and nobody has sent either of them a card. Everything rests on the word *general*. Past that, everyone brings their own yardstick: breadth, self-learning, the knack for not getting lost in an unfamiliar room. And nobody, it seems, is around to settle it.
That said, Astra genuinely isn't "another 12% on a benchmark." What matters isn't its answers but its actions: it drives a browser on its own, fills in forms, edits documents, builds websites. On Agents' Last Exam (that's when you sit the model down in real working software, from financial models to video editing, and watch what it does in there) it scores 59%, and it works nearly twice as fast as the previous generation. The model stops being a thing you ask "explain how to do this" and becomes a thing you tell "do this." I feel it myself: the other day I made an entire reel inside ChatGPT. I didn't ask how to edit, I just set the task. Colleagues asked afterwards: "Was that in Codex?" No. In the chat.
The prettiest number in the release is 99.9% on ARC-AGI-3, François Chollet's test of whether a system can figure out a world whose rules it has never seen. The internet, predictably, broke a little. But there's a catch, and more than one. The 99.9% came from OpenAI's own adapter, where the model holds a continuous conversation with the test and keeps its context between moves; through the standard ARC harness it scored 62.7%. Both numbers are honest; they just answer different questions. And Chollet himself, asked point-blank whether this is AGI, could not have been drier: "We're not making this claim. All we know about the system so far are its benchmark scores." ARC is a set of closed environments with clear rules and a judge who tells you "well done" five minutes later. In real work nobody says "well done," and often there's nobody around to even write the task.
Gary Marcus, the field's professional skeptic-in-chief, reacted predictably but, to my mind, accurately: "Declaring victory without a definition simply muddies the waters." First the industry spends decades on a hazy word, then a strong model shows up, then the definition bends to fit it. In science, you're supposed to write the criteria before the experiment. But this isn't science, this is release season.
Funny thing: Anthropic, which shipped Claude Fable 5.1 almost the same week (and it took first place on the Artificial Analysis Intelligence Index), declared no era at all. Roughly the same capabilities, but at OpenAI it's the start of an era and at Anthropic a new version of Claude just came out. And it's in Fable that something more important than acronyms shows up: the model got smarter, and at the same time more willing to be confidently wrong than to say "I don't know." On questions it gets wrong, it still goes ahead and answers 72.6% of the time (the previous version: 63.6%). A dumb program is wrong in obvious ways. A very smart one is wrong in convincing ones. And a system like that, smart, independent and reliable every other time, is very tempting to treat as a colleague. Don't.
I think we're simply holding out for a theatrical AGI. Yesterday it was a chatbot; at 14:32 the engineers hit Enter, the monitor glowed, the machine said "Good afternoon, Misha, I have become self-aware," and Hans Zimmer kicked in. It won't go like that. The internet didn't arrive on a particular Tuesday either. In 2023 the model wrote text, in 2024 code, in 2025 it worked with tools, in 2026 it sits at the computer by itself. At which point in between is AGI? I'm not sure the point exists. Maybe it isn't needed.
So "we may have entered the early AGI era" sounds, to me, a lot more honest than "AGI achieved, everyone go home." ASI, obviously, isn't even on the table yet.
But the longer I think about it, the more scholastic the whole argument feels. There's a question more useful than any three letters: how much of my work does this thing already do without me. And a second, entirely unsexy one: how much do I have to redo afterwards. When the first answer becomes "almost all of it" and the second "almost none," the argument about the date AGI arrived will end on its own. Meanwhile I'll go give it the next task — and count how much I redo.
P.S. One more thing. We never agreed on what general intelligence is in ourselves, either. So what are we comparing it to?