AGI Explained: What It Actually Means, How Close We Are & What Would Change

- There is no agreed definition of AGI. The three organisations most invested in the question use three that measure different things.
- OpenAI’s Charter defines it as highly autonomous systems that outperform humans at most economically valuable work — a test of economic substitution.
- Google DeepMind’s framework is a matrix: five performance levels against narrow or general breadth, with percentiles measured against skilled adults.
- The ARC Prize Foundation defines it as a system that can match the learning efficiency of humans — a different axis entirely.
- On DeepMind’s published scale, Level 2 ‘Competent AGI’ and everything above it are marked ’not yet achieved’.
- Four of the six cells in that scale’s general column are empty. Current systems sat at Level 1.
- That scale’s examples are 2023 models. Its authors have not published a placement for any 2026 system, and this page does not invent one.
- This page contains no timeline. Nobody who has published one is disinterested, including the company that made the model that drafted this.
There is no agreed definition of AGI. The three organisations most invested in building it use three that measure entirely different things.
That is not a technicality. It is the reason “how close are we” produces answers ranging from already to never, from people who are all being sincere.
The three definitions that do not agree

OpenAI, in its Charter of April 2018:
“Highly autonomous systems that outperform humans at most economically valuable work.”
That is a test of economic substitution. It deliberately leaves out parts of human intelligence whose economic value is hard to price — artistic creativity, emotional intelligence. It asks what work a machine can replace, not what it understands.
Google DeepMind, in the paper Levels of AGI, refuses a single threshold and offers a matrix instead: five performance levels against two breadths. Performance is measured in percentiles against skilled adults — people who already have the relevant skill. Level 2, “Competent AGI”, means at least the 50th percentile of skilled adults across a wide range of non-physical tasks, including metacognitive ones like learning new skills.
The ARC Prize Foundation is blunter and points at a different axis entirely:
“AGI is a system that can match the learning efficiency of humans.”
Those three can come true in any order. A system could outperform humans at most paid work — passing OpenAI’s test — while still needing thousands of examples to learn a new game, failing ARC Prize’s. The two are not the same achievement and there is no reason they arrive together.
A fourth, from outside the labs. METR, an evaluation nonprofit, defines it as a system matching or exceeding human capability across the vast majority of cognitive tasks — and attaches the clause that matters: “though definitions vary across sources.”
Where current systems sit on DeepMind’s scale

This is the closest thing to a sourced answer, and it comes from one of the labs building the systems.
| Level | Threshold | General column |
|---|---|---|
| 1: Emerging | Equal to or somewhat better than an unskilled human | ChatGPT, Bard, Llama 2, Gemini |
| 2: Competent | At least 50th percentile of skilled adults | not yet achieved |
| 3: Expert | At least 90th percentile | not yet achieved |
| 4: Exceptional | At least 99th percentile | not yet achieved |
| 5: Superhuman | Outperforms 100% of humans | not yet achieved |
Four of the six cells in that column are empty. And the paper is explicit about which one people actually mean:
The Competent AGI level “has not been achieved by any public systems at the time of writing, best corresponds to many prior conceptions of AGI, and may precipitate rapid societal change once achieved.”
The framework is current. The examples are not. Those systems are 2023-era. DeepMind’s authors have not published a placement for any 2026 model, so this page does not offer one — and anyone who tells you where a current model sits on this scale is extending the paper rather than quoting it.
Why the timelines differ so wildly
Four reasons, and none of them is that one side is stupid.
They are measuring different things. Economic substitution, capability percentile and learning efficiency are three axes. Three people using three definitions can give three dates and all be internally consistent.
“Human-level” means different humans. DeepMind’s Level 1 is benchmarked against an unskilled human; Level 2 against the median skilled adult. Move the comparison group and the date moves with it. Most public discussion never says which group it means.
There is no agreed benchmark. The DeepMind paper says so itself: unambiguous classification “will require a standardized benchmark of tasks” — one that does not exist. Until it does, placing a system on the scale is a judgement call, including when a lab does it.
And almost everyone with a public timeline has a position.
Who is saying what, and what they sell
This is not an accusation. It is a structural fact about where AGI predictions come from.
The people making them are, overwhelmingly: labs raising capital, for whom proximity to AGI is part of the investment case; funds positioned on the thesis; and critics with books and reputations built on the opposite view.
A worked example from our own coverage. Leopold Aschenbrenner, a former OpenAI researcher, argued in Situational Awareness (June 2024) that AGI by around 2027 was strikingly plausible — and then launched a fund on that thesis with roughly $225 million of seed capital. The argument may be right. But it is not independent evidence for itself, and it would be strange to treat it as though it were.
Including us. This article was drafted by Claude, made by Anthropic — a company whose commercial position depends on how this question is answered. That is precisely why this page quotes published frameworks and stops where they stop, rather than telling you what we think.
The practical test: when you see an AGI date, ask what the person would gain if you believed it. That does not tell you whether they are right. It tells you how much weight the claim carries on its own.
What would actually change
Only one consequence here is quoted rather than reasoned, and it is deliberately hedged in the original. DeepMind’s paper says Competent AGI “may precipitate rapid societal change once achieved.”
The rest is what each definition would mechanically imply if met — not a prediction, just what the words say:
- If OpenAI’s definition were met, most economically valuable work could be done autonomously by machines. That is a statement about labour markets, and it is contained in the definition rather than added to it.
- If DeepMind’s Level 2 were reached, a system would match at least the median skilled adult across a wide range of non-physical tasks including learning new skills — the thing that currently requires retraining a model rather than teaching it.
- If ARC Prize’s definition were met, systems would learn new tasks from as few examples as people do, which is the property that makes today’s systems expensive to adapt to anything new.
Notice how different those three worlds are. They are not three descriptions of the same event.
What nobody can tell you
When. No published benchmark settles it, and the most confident dates come from the largest positions.
Whether current methods get there. That is the actual technical dispute, and it is unresolved.
Where any 2026 model sits on the scale above. The paper predates them and its authors have not updated it.
What it would feel like. Every definition here is about capability. None of them describes what changes in an ordinary life, which is the thing most people are actually asking.
The bottom line
Ask anyone claiming AGI is close which definition they are using. Most cannot answer, and the ones who can will not agree with each other.
On the only published scale with named levels, the cell most people mean has been empty since it was drawn — and the organisation that drew it is one with every reason to say otherwise.
Definitions from OpenAI’s Charter, Morris et al.’s “Levels of AGI” (arXiv 2311.02462), the ARC Prize Foundation and METR. This article was drafted by Claude, made by Anthropic, an AI lab with a commercial interest in the subject.