What is AGI? Depends who's selling it

Tags
agiaillm

Everyone is racing toward AGI. Nobody agrees on what it is. A short history of a word that keeps changing, and what I think it actually means.

It's September, and two things happened in the space of three weeks.

On the 3rd, OpenAI shipped GPT-6 Astra, and Greg Brockman said it may one day be seen as the arrival of AGI. On the 22nd, Anthropic shipped Claude Opus 5.5, which beats Astra on most of the benchmarks people use to argue about these things.

So which one is AGI? Both? Neither? Did we cross the line on the 3rd, or did the line just move again on the 22nd?

If the finish line had a definition, we'd know. It doesn't. 🤔

AGI has no definition. It has a history. And that history tells you more about where we are than any benchmark card.

What is AGI? Five definitions that don't agree

Ask five serious people what AGI means and you'll get five answers that sound alike and measure completely different things.

  • OpenAI's charter talks about "highly autonomous systems that outperform humans at most economically valuable work." That one is about jobs.
  • Google DeepMind's "Levels of AGI" paper (2023) lays out a grid of performance and generality, from "emerging" to "superhuman." Capabilities, in other words.
  • The Microsoft and OpenAI contract, as reported at the end of 2024, said AGI arrives when OpenAI's systems can generate around one hundred billion dollars in profit. That's a definition about money. The October 2025 restructuring handed the call to an expert panel, and in April 2026 the clause was dropped altogether, which tells you how much weight the word could actually carry in a contract.
  • Engineers tend to say: anything a human can do at a computer, the model can do.
  • Mine is a system that can solve any problem in mathematics, science or code, and then put the solution in place, at a level no human can match. It's about outcomes, and you can check it field by field.

Put them side by side and they barely overlap. A system could pass one and fail the other four.

That's what happens to a word that means everything. It ends up meaning whatever the person saying it needs. For a lab, it's a mission. For an investor, it's a valuation. For a contract, it was an exit clause. For a critic, it's a line that is always one more year away.

Where the term AGI came from

AI started out as AGI. It just didn't need the "G."

The 1956 Dartmouth workshop that named the field aimed to simulate "every aspect of learning or any other feature of intelligence." Then came two AI winters, and the field survived by going narrow: chess engines, speech recognition, spam filters.

The phrase "artificial general intelligence" first appeared in 1997, in a paper by Mark Gubrud that almost nobody read. It stuck around 2002, when Shane Legg, later a DeepMind co-founder, suggested it to Ben Goertzel as a book title. Artificial General Intelligence came out in 2007, and a small community finally had a name.

So the "G" wasn't a new idea. It was added because the field had forgotten what it originally wanted.

How the meaning moved

Since then the word has changed jobs at least three times, and each time someone new had a reason to use it.

In the 2000s it was a research goal, and a fringe one. Nobody gained much from saying it. The conferences were small, and admitting you worked on it in a serious lab didn't do much for your career. It meant a machine that could learn anything a person could learn.

In the 2010s it became a mission statement, and the labs were the ones who gained. DeepMind was founded in 2010 to "solve intelligence." OpenAI came along in 2015, and its charter later promised that AGI "benefits all of humanity." A mission that big recruits the best researchers and justifies the biggest cheques. The word went from the fringe into the founding documents of the best-funded labs in the world.

After ChatGPT it turned into a product claim, and now the sellers gain. AGI became something you launch, raise money on and write into contracts. Every major release arrives with the question attached, whether the lab asks it or not, and the valuation moves with the answer.

And the whole time, the goalposts kept moving.

Chess was supposed to require real intelligence, until Deep Blue beat Kasparov in 1997. Then it was "just search." Go was supposed to be different, until AlphaGo in 2016. Then it was "just pattern matching." The Turing test held its place for seventy years, and people quietly stopped mentioning it once language models walked through it. Gold at the International Mathematical Olympiad was the next wall. That fell in 2025.

Every time a machine does the thing, the thing stops counting. I don't think that's dishonesty. It's what happens when you try to define intelligence by listing examples of it.

Where AI models stand in 2026

Forget the word for a moment and look at what the models are doing this year. I wrote last month that the AGI race has two scoreboards. Benchmarks are the one that gets reported, and "AGI" is the headline people write on top of it. So here is that scoreboard, plus what doesn't fit on it.

Start with mathematics. In May, an OpenAI model found a counterexample to the unit distance conjecture, a problem in discrete geometry that Paul Erdős posed in 1946. Eighty years of mathematicians, and the answer came from a machine. OpenAI followed with ten more results on long-standing open problems across geometry, coding theory, complexity and cryptography, and Quanta ran a whole piece on why the Erdős problems keep falling to AI.

This month OpenAI went further, with a result on Navier–Stokes, one of the Millennium Prize Problems. Quanta ran it under the headline AI has solved one of them. The proof has been checked in Lean, a system that verifies mathematics line by line. Mathematicians are still arguing over how much credit belongs to earlier human work and whether it meets the prize criteria, which is exactly what they should be doing. A year ago nobody would have taken the sentence seriously enough to argue about it.

Research math

FrontierMath Tier 4 is research-level mathematics, problems that can take a specialist days. In April, GPT-5.5 solved about a third of them. By September, Astra solves almost all of them (lab-reported). For context, a year ago the best model solved 13% of an earlier version of the set. Anthropic has not published an Opus 5.5 score. Source: Epoch AI, OpenAI, September 2026.

Humanity's Last Exam was designed to be the final exam for AI. When it launched, the best models scored in single digits without tools, and about 27% with them. Opus 5.5 now scores 67.7% with tools.

Humanity's Last Exam

Built in January 2025 from questions experts wrote to stump the best models. Twenty months later, two thirds of it is answered. Source: Center for AI Safety & Scale AI, lab reports, September 2026.

The senses caught up too. The assistants we deploy on Ardaven take a customer's voice note, a photo of a damaged product or a screenshot of an error, and answer all three from the same business knowledge. Not long ago that took separate systems, glued together badly.

And it's uneven, which is the part the headlines miss. Astra leads Opus 5.5 on Terminal-Bench-Science, 64.6% to 58.7%. Opus 5.5 leads Astra on Humanity's Last Exam by more than ten points with tools, and about seven without. The same model that finds a counterexample Erdős missed can still get confused by a spatial puzzle a child solves in seconds.

People call this jagged intelligence: superhuman in some directions, strangely weak in others. It breaks any definition that treats intelligence as one number you can rank.

But intelligence isn't one number. It never was for us either. A great mathematician can be hopeless at reading a room. A brilliant surgeon might not be able to write a decent paragraph. We never asked people to be equally good at everything before calling them intelligent.

The physical world is where we're behind

There is one area where the gap is real, and it's the physical world.

Understanding a room from a picture is one thing. Walking into that room, opening a stuck drawer and folding the shirt inside is another. The models can describe the drawer beautifully. Robots still struggle to open it.

That matters, because a lot of human work happens with hands, not keyboards.

Even so, the last two years here have been huge. Robotics moved from hand-coded controllers to foundation models. Vision-language-action models let a robot take an instruction in plain language and turn it into movement, and the newest systems add world models, so a robot can imagine what happens next before it moves. NVIDIA's GR00T and Cosmos Reason are open models built for exactly this. For the first time, robot training data is starting to show the kind of scaling behaviour that made language models take off.

Still, 2026 is the year robots are proving themselves in narrow industrial and logistics work. A general-purpose robot in your kitchen is a few years off, because reliable autonomy in messy, unstructured places is hard.

The mind got there first. The body is catching up, faster than it looked two years ago. 🦾

What AGI means to me

So here is where I land.

My definition, again: a system that can solve any problem in mathematics, science or code, and put the solution in place, at a level no human can match. I set the bar high on purpose. Better than any person, and the solution has to ship to production environments, because finding it is only half the job.

No model clears that bar today. But I don't expect to wake up one morning and find it cleared. For me that bar is the far end of a range, and we're already inside the near end.

The near end has a simple test. For years, every time a model did something impressive, we added a qualifier: "for an AI, that's good." Good code, for an AI. A clever proof, for an AI. That qualifier was doing a lot of work. It meant we were grading on a curve.

You'll know a field has entered the range when you stop saying "for an AI, that's good." When the work is simply good, and you review it the way you'd review a colleague's, and nothing about the review changes because a model produced it.

Every definition above tries to pick a single moment, like the day a model beats humans at most jobs or the day it makes a hundred billion dollars. That moment won't come as a clean announcement, because the capability doesn't arrive all at once. It shows up field by field, unevenly, one benchmark and one Erdős problem at a time. Mathematics is well inside the range. Coding is inside it. Scientific reasoning is close. The physical world is at the edge and moving in.

When a model finds a counterexample to an eighty-year-old conjecture, nobody says "for an AI." It's a result, and one no human had managed. That's the range, entered in one corner of one field. The far end is still ahead. In other fields, especially anything with hands, I still say the qualifier every day.

From where I sit, AGI looks like a dial, turning at different speeds in different fields, and already a long way from zero.

What to do with this on Monday

You don't need to settle the definition to use it.

  • Run the "for an AI" test on your own work. For each kind of task you hand to a model, ask whether you still grade it on a curve. At LayerX we dropped the qualifier for first drafts of code months ago: agents write most of it, and we review it like any colleague's pull request. We haven't dropped it for architecture. That line tells us where to trust and where to steer.
  • When a vendor says AGI, ask which definition, and what it gets them. Jobs, capabilities, money or outcomes. The answer tells you more about the pitch than about the model.
  • Measure field by field, not with one number. A model that leads on maths may trail on science tasks. Your own evals on your own work are the only scoreboard that counts.
  • Spend your effort on the layer around the model. A field entering the range doesn't make a model safe to deploy. So Ardaven, our agent platform, answers only from your own documents, logs every message with the model, tokens and tool calls behind it, and gives admins a kill switch on any conversation. The model gets smarter every month. The layer around it is what lets you put it in front of customers.

The intelligence is arriving. The engineering, auditability, orchestration and the governance around it are still our job.

Related posts