AGI means whatever we want it to mean


Greg Brockman closed the GPT-6 Astra briefing with “Welcome to the AGI era.” Once again, Brockman has no clue what he’s talking about.

In Empire of AI, Karen Hao quotes Brockman on the subject of AGI. Regarding climate change: “It’s a super-complex problem. How are you even supposed to solve it?” When discussing medicine: “Look at how important health care is in the US as a political issue these days. How do we actually get better treatment for people at lower cost?”

Brockman may be the president of OpenAI, but he is not a serious person on the subject of AGI. We haven’t failed on climate change and health care because we lack the right model weights. We lack the political will to change the systems we live in. Someone who can’t clearly see the difference between a large language model and an economic-political system should have no say in either.

We know what the climate crisis requires: renewable energy, a carbon tax with penalties big enough to change corporate behavior, and a global commitment to reduce consumption.

We know what the health care crisis requires: universal coverage, early intervention, price transparency, and a system that prioritizes patient outcomes over profit.

These are not problems that can be solved by a model, no matter what generation. The model will do whatever humans tell it to do. If corporations choose to continue to maximize profits at the expense of human life, the model will assist.

Nvidia CEO Jensen Huang congratulated OpenAI on the launch of Astra. “From ChatGPT to o1 to Astra in 4 years,” he wrote. “AGI has arrived.” In the same post, he noted that Astra trained on more than 100,000 Nvidia Grace Blackwell NVLink72 systems, and that another 400,000 GPUs come online next.

The man certifying AGI sells the GPUs. This isn’t his first certification, either. He declared it about six months ago, on an earlier model.

Moving the goalposts

OpenAI’s charter, written in 2018, defines AGI as “highly autonomous systems that outperform humans at most economically valuable work.” That’s not a technical benchmark. It’s the economics of labor. Altman reiterated the same standard himself in a 2025 blog post, writing that OpenAI can build AGI “as we have traditionally understood it.”

Fair enough. We have definitions, no matter how flawed they may be.

Except that isn’t what happened last week. Brockman, on the record, said, “I do leave it up to the reader to decide for themselves if this qualifies for them.” That is not a company reaching its objective. That’s OpenAI standing next to a line it drew itself, claiming to have crossed it, then telling you to be the referee.

Google DeepMind’s own AGI taxonomy identifies ten cognitive abilities and a three-stage evaluation framework, including community benchmarks. None of its metrics are economic.

62.7% vs. 99.9%

Here’s the part that should have been the headline instead of the AGI quote.

Astra’s marquee result was 99.9% on ARC-AGI-3. OpenAI ran Astra through its own provider adapter harness, which preserves the model’s reasoning state between requests and compacts long runs.

The nonprofit ARC Prize ran the same model through its own standard harness, which is open source and reproducible. Under that harness, Astra scored 62.7%.

A 37-point swing. Same weights, different scaffolding. The harness you can audit says 62.7%. The one you can’t says 99.9%.

ARC Prize didn’t bury this. They report both scores separately on their leaderboard, which is the polite, technical way of saying: don’t cite the 99.9% without the asterisk. They went further and refused the conclusion OpenAI drew from their own benchmark. ARC stated it is “not claiming that it is AGI.” Co-founder Mike Knoop wrote that “we lack evidence to call this AGI yet.”

$26,098 vs. $12.78

By ARC Prize’s own accounting, the standard-harness run cost $26,098 to score 62.7%. The provider adapter run cost $18,817 to reach 99.9%. The human baseline came from about 500 members of the general public, recruited across a range of ages, incomes, and occupations, with no puzzle-solving pedigree. They got $115 for a 90-minute session, plus $5 for each environment solved. Participants attempted roughly nine games per session. ARC puts that at about $12.78 per attempted game.

$12.78 against $26,098. Humans can solve 100% of the environments. ARC only includes an environment if it clears an “easy for humans” bar.

Billions of dollars to build it and tens of thousands to run it, to land under two-thirds on a test built to be easy for a stranger off the street. That’s not a demonstration of general intelligence. That’s a demonstration of what you can do with an unlimited budget and no obligation to be efficient.

By any measure, that fails to “outperform humans at economically valuable work.” Economically valuable work has a budget.

Lying liars

You don’t need to lie to say something is good.

None of this means Astra is a bad model. The action-efficiency numbers on ARC-AGI-3 are real, and computer use is measurably better than its predecessor. That’s a legitimate generational step.

It’s another model. Better? Sure. But Astra is not AGI. Not by DeepMind’s definition, and not by OpenAI’s own.