A system that matches or exceeds human performance across the full range of cognitive tasks rather than one narrow slice. There is no agreed definition and no accepted test, so progress is argued through benchmarks that stand in for the term.
AGI names a system that handles the full range of cognitive work a person can handle, rather than performing well on one narrow slice of it. The systems in production today sit on the narrow side of that line: they are strong across a wide span of language tasks and weak or absent on others, and they do not carry what they learn in one session into the next.
There is no agreed definition of the term and no accepted test for it. Every claim that a system has reached AGI, or is close to it, rests on a definition chosen by whoever is making the claim, so the first question to ask of any such claim is which definition is in use.
Benchmarks stand in for the missing definition. ARC-AGI, GPQA, FrontierMath, SWE-bench and Humanity's Last Exam each measure one slice, and each follows the same arc: a score climbs, the benchmark saturates, it is retired, and a harder one replaces it. That cycle is the reason benchmark results move the conversation without settling it.
Two problems sit underneath. Test items leak into training data, so a model can score well on questions it has effectively already seen. And a benchmark measures the task it contains, not the general capacity the task was chosen to represent, which is the same gap that separates a high exam mark from competence at a job.
One camp holds that scale is the path: more parameters, more compute, more data, and the remaining gaps close on the current trajectory. The other holds that specific ingredients are missing, and names them: continual learning, so a system improves while it runs rather than starting every task cold; world models, so a system predicts consequences rather than text; and sample efficiency, so a system learns from a handful of examples rather than a corpus.
Our article 2037. AGI. China probably. sets out where we stand on this. Continual learning is the gating requirement, the world model needed to support it is light rather than exhaustive, and the release cadence and cost discipline of the Chinese open-weight labs make them the group to watch.
Nothing in current AI visibility work depends on the term. The systems that decide whether a brand is cited or recommended are large language models with retrieval attached, and AI agents running inside an agentic harness. They are narrow in the sense that matters here, and they will behave the same way tomorrow whether or not the AGI label is claimed by anyone.
The shift already underway is delegation. A growing share of tasks now runs through an assistant or an agent, so fewer people reach a page directly, and the surface being competed for is the answer rather than the result listing. That is a change in how systems are deployed, and it arrives on its own schedule regardless of where the definition of AGI lands.