Research
What researchers call AGI
There is no single definition. There is a cluster. Aether treats the cluster as a specification: if a lattice cannot move a system on these axes, scaffolding is not enough. If it can, we still have not built a mind — we have built a better prosthesis.
- 012025
Hendrycks, Song, Szegedy, Bengio, Schmidt, Marcus, Tegmark et al.
AGI is an AI that can match or exceed the cognitive versatility and proficiency of a well-educated adult. Operationalized as ten equally weighted CHC broad abilities.
A score, not a vibe. GPT-4 27%, GPT-5 57%. The remaining hole is foundational machinery — especially long-term memory storage at 0.
This is the definition Aether scores against. Scaffolding is the bet that MS, MR, WM, and R can be raised as prostheses.
- 022023–24
Morris et al., Google DeepMind — Levels of AGI
Performance (Emerging / Competent / Expert / Virtuoso / Superhuman) crossed with generality (Narrow vs General). Autonomy is a separate axis.
Frontier chat models sit at Emerging–Competent AGI: they cover a wide range of tasks at or below a skilled adult, not Expert across the board.
Aether's battery is a Competent-AGI probe. Expert would require the same lattice to hold over hours of tool-using work, which this product can attempt but not certify.
- 032007
Shane Legg & Marcus Hutter
Intelligence is an agent's ability to achieve goals in a wide range of environments.
Generality is environmental, not benchmark-shaped. A language model in a chat box is one environment.
The chamber is one environment. Tools, memory, and goals widen it. Embodiment and the physical world are still missing.
- 042019–25
François Chollet — On the Measure of Intelligence / ARC-AGI
Intelligence is skill-acquisition efficiency: how little experience you need to master a novel task. ARC tests fluid abstraction, not memorized skill.
o3-class systems jumped ARC-AGI-1 via test-time search, then fell on ARC-AGI-2. Search can look like fluid intelligence until the distribution shifts.
Planner + debate is test-time search. Aether will not pretend a high battery score is Chollet-AGI. The reasoning probe is deliberately novel.
- 052018–25
OpenAI Charter (and later internal levels)
Highly autonomous systems that outperform humans at most economically valuable work. Informal ladder: chatbots, reasoners, agents, innovators, organizations.
An economic definition. It can be met by a jagged specialist swarm without CHC completeness (no genuine MS, no audition).
Aether is not a firm. Goal stack + tools are a toy agent level. Economic AGI is out of scope and honestly labeled so.
- 062026
Burnell, Morris, Botvinick, Legg et al. — Measuring Progress Toward AGI
Place systems on the human performance distribution per cognitive faculty. Coverage is still missing for metacognition, attention, learning, and social cognition.
Jagged profiles are the expected shape. A mean AGI% hides valleys.
Observatory shows the jagged profile on purpose. Metacognition is an explicit layer because DeepMind flagged it as a coverage hole.
- 072025–26
METR — time-horizon / autonomy evaluations
The relevant number is how long a model can pursue a task unattended before humans have to intervene.
Harness quality (recovery, checkpoints, verification) moves this number as much as weights. SWE-bench Pro analyses in 2026 found harness swaps beating many model swaps.
LangGraph-style state, memory, and reflexion exist to extend horizon. Aether's chamber is short-horizon by design; the battery records whether a fact survives a later turn.
- 082000s–
Ben Goertzel
AGI as a system with the capacity for general intelligence comparable to humans, including learning, transfer, and autonomous goal pursuit — often via a cognitive architecture, not a single net.
The original 'AGI' community already assumed scaffolding: memory, perception, reasoning modules around a core.
Aether is a small cognitive architecture around Grok. That is the Goertzel-shaped bet: generality from the system, not only from the model.