LIVE
Publish Flash items in Admin to fill the ticker
The Roman Road of AI in the New Era
Sign InSubscribe ProAdmin
FrontierFREE

AGI Roadmap Part 1: History, Frontier Paths, and the 2026 Present

|

Comprehensive AGI roadmap Part 1: history from Dartmouth to agents, eight Frontier paths as a bottleneck stack, and a 2026 capability snapshot. Not a countdown.

AGI Roadmap Part 1: History, Frontier Paths, and the 2026 Present

Evidence note: Frontier Desk synthesis from the site’s Frontier path taxonomy, public lab research pages, widely cited papers, and standard historical accounts as of 2026-08-23. This is a map of definitions, research waves, and interacting bets—not a timeline, scoreboard, or date for AGI. We do not invent unpublished lab scores or treat rumor names as facts.

Quick answer

An AGI roadmap is useful only if it refuses to be a countdown. Part 1 of this series treats artificial general intelligence as a bundle of competences—transfer under shift, long-horizon reliability, tool-using agency, checkable reasoning—not a single threshold a lab will announce. The historical record from Dartmouth (1956) through two winters, expert systems, statistical machine learning, the 2012–19 deep-learning wave, the 2020–23 scale wave, and the 2024–26 reasoning-and-agents wave shows the same pattern: each era solved a narrower problem than its press cycle claimed, then left a residue that later methods had to attack. The Frontier desk therefore reads eight research paths as a bottleneck stack, not as rival churches: Scaling / LLMs still supplies the general approximator; world models try to buy compact prediction and planning; RL & agents close the loop from prediction to policy; embodied AI buys contact-rich grounding; neuro-symbolic methods buy structure and checkers; multi-agent systems buy specialization (and extra failure surface); AI for science buys verifiers and synthetic curricula; neuromorphic / hardware sets the energy floor under every other loop. The 2026 present is a capability snapshot with large productized islands and equally large gaps. Software, hardware, and scenario futures belong in Part 2. This article stops at the map of how we got here and which constraints are live.

Key takeaways

  • AGI is a spectrum of competences, not a finish line. Fluency, exam scores, and demo videos are not the same object as reliable transfer, long-horizon action, or physical grounding.
  • History is a sequence of over-claimed winters and under-claimed residues. Expert systems, statistical ML, and deep learning each left a bottleneck the next wave inherited. Scale did the same.
  • The eight Frontier paths are layers, not a menu. Read them as which constraint is binding this year. The shorter connector, how frontier paths connect, is the coupling diagram; this Part 1 is the history and path deep-dive.
  • Scaling remains the engine and a contested complete theory. Diminishing generic-pretrain returns do not mean “scale is dead.” They mean other axes—data quality, post-training, test-time compute—start to bind. See where scaling returns diminish.
  • 2026 productization outruns 2026 generality. Assistants, coding loops, and short-horizon tools shipped. Long-horizon agency, cheap inner-loop physics, and checkable open-world action did not. Those gaps are the brief for Part 2.
  • Ignore countdown threads until a bottleneck is named. Use this map as a routing tool for research, product, and capital—not as a prophecy.

Who this is for — and who should skip

Read this if you allocate research, product bets, or attention across AGI-adjacent work and need a shared vocabulary: founders deciding where not to compete with labs; researchers placing a paper on a public road rather than a slogan; operators translating announcements into the 2026 model stack; readers who want history and definitions before software/hardware futures. Pair lab headlines with lab watch 2026 rather than with social-media arrival dates.

Skip this if you need a ship checklist for chat, retrieval, or a single tool loop. Start with the chatbot-to-agent map, then come back. Skip also if you want a year, a lab, or a “who is winning AGI” table. This desk does not publish those. Part 1 is a route synthesis with a long historical section. It is not a tutorial and not a forecast.

How Part 1 relates to Part 2 and the shorter connector

This series is three artifacts on purpose. The short connector, The Road to AGI: How the Frontier Paths Connect, is the one-sitting map: eight paths, a coupling list, a routing table. Read it when you already know the vocabulary and want the stack picture. Part 1—this article—is the long brief: definitions, what AGI is not, measurement axes, the post-1950s history that produced the current bets, a deep dive on each path, and a 2026 snapshot of what is actually productized versus what remains a research residue. Part 2, software, hardware, and futures, takes the residue as given and asks what software stacks, energy and chip constraints, and scenario families look like if you refuse prophecy. Do not collapse the three into one countdown essay. The connector is the diagram. Part 1 is the archive and the present. Part 2 is the fork, not the date.

A reader who only has time for one piece should take the connector. A reader who will argue about “is this AGI?” should take Part 1. A reader who will budget inference, robots, or multi-year research should take both parts in order. The Frontier path page remains the live taxonomy; climate tags there are editorial descriptors, not probabilities.

Why a roadmap is not a countdown

Public AGI talk collapses into two failure modes. The first is a countdown: a lab, a year, a threshold, a leaked name. The second is a silo war: scaling is dead, world models will replace transformers, embodiment is mandatory, embodiment is a detour, multi-agent is civilization, multi-agent is an org chart with extra latency. Both failure modes are easy to publish and hard to falsify. A countdown cannot be checked until the date passes, at which point the goalposts move. A silo war treats complementary constraints as religions.

The Frontier premise is blunter. General competence is a bundle of bottlenecks. Prediction quality, planning horizon, action interface, systematic structure, coordination cost, verification, and watts do not yield to a single architecture press cycle. Labs already hire, publish, and spend as if that were true—even when keynotes pick a mascot path. Connecting the paths is more honest than ranking them. Ranking them is how you get a prophecy. Connecting them is how you get a research agenda you can update when a paper actually moves a constraint.

This article therefore uses “roadmap” in the civil-engineering sense: a survey of terrain already crossed, a legend for the roads still under construction, and a list of bridges that do not exist yet. It is not a Gantt chart with AGI in the last cell. If you came for a date, the honest output is that this desk will not give you one. If you came for a map, the rest of the article is the legend.

I. Definitions: spectra, exclusions, and measurement axes

Most fights about AGI are fights about the noun. If the word means “a system that can do any economically valuable cognitive task at or above a median knowledge worker,” you are talking about a labor-market object. If it means “a system that can acquire new skills with the sample efficiency of a human child,” you are talking about Chollet-style skill-acquisition efficiency. If it means “a single set of weights that matches or exceeds the best specialist on every benchmark we currently run,” you are talking about a leaderboard object that can be gamed. If it means “an autonomous agent that pursues open-ended goals in the physical world,” you are talking about robotics plus governance. These are not the same research program. A roadmap that pretends they are will smuggle a definition into a forecast.

AGI as a spectrum, not a binary

This desk treats AGI as a spectrum of competences that can be present in different degrees, in different domains, at different costs. A useful spectrum has at least four axes that move somewhat independently:

  • Breadth. How many domains can the system enter without a new architecture? A coding-strong model that collapses on spatial planning is narrow in a way a 1990s expert system was also narrow—just at a different scale of fluency.
  • Transfer. When the surface form of a task changes, does competence survive? Exam-style benchmarks often measure retrieval of a training-adjacent template. Transfer measures whether the template was a skill or a souvenir.
  • Horizon. How many sequential decisions can the system take before error compounds into uselessness? Next-token chat has horizon one in the product sense. A month-long research agent has a horizon that current traces rarely survive.
  • Agency under constraint. Can the system change external state—files, tickets, instruments, actuators—while remaining inside a budget, a permission set, and a stop condition? This is the product definition of an agent in 2026, and it is not the same as “sounds autonomous in a demo.”

A system can be broad and shallow (fluent everywhere, reliable almost nowhere), deep and narrow (superhuman in protein geometry, clumsy in a help desk), or long-horizon in a simulator and short-horizon in a messy institution. Calling any one of those “AGI” or “not AGI” is a branding decision. Mapping where it sits on the spectrum is a research decision.

Adjacent vocabulary that this series will keep distinct: ANI (artificial narrow intelligence) is competence locked to a task family. ASI (artificial superintelligence) is a speculative regime beyond human specialist competence across many domains; this desk does not treat it as an operational 2026 category. Foundation model is an industrial term for a large pretrained approximator that is adapted downstream; it is a stack layer, not a claim about generality. Agent in the 2026 product sense is a control loop with tools. World model is an internal predictor used for imagination or planning. None of these words should be used as synonyms for AGI. They are components or marketing umbrellas.

What is not AGI (and why the exclusion list matters)

Exclusions are more useful than slogans because they stop category errors before they become procurement errors. The following are not AGI on this desk, even when they are impressive, valuable, or necessary ingredients:

What is not AGI — common substitutions and why they fail (desk heuristic, 2026)
Object people call “AGI” What it actually is Why the substitution fails
A high score on a public exam-style benchmark A measurement on a contaminated or saturated probe Does not establish transfer, horizon, or agency
A viral robot or computer-use demo A curated trajectory under hidden resets Video is cheaper than held-out reliability
A chatbot that “feels general” A strong next-token approximator plus product UX Fluency is not a policy; tone is not a world model
A reasoning mode that spends more tokens Test-time compute on a base model Search can raise competence without raising generality
A multi-agent org chart in a graph UI Coordination overhead on top of one loop More roles do not create a source of truth
A science result in a closed domain A verifier-rich specialist system A real reward is not open-world common sense
An open-weight drop that matches last year’s API An access-mode event Distribution is not a new competence class
A lab keynote that uses the word AGI A narrative and recruiting instrument Slogans are not eval harnesses

The exclusion list is not cynicism. High scores, demos, chat, reasoning modes, multi-agent graphs, science results, and open weights are all real. They are just different objects. A roadmap that lets them collapse into one noun will keep announcing arrival and keep being surprised by the next residue.

Two further exclusions that come up in comments. First, human-likeness is not the criterion. A system can be general while being unlike a person in memory, sensorium, and cost. Insisting on a humanoid body or a human childhood as the definition smuggles embodiment and developmental psychology into a claim about competence. Those may be necessary for some competences. They are not a definition. Second, consciousness, sentience, and moral patienthood are not this article’s subject. They are philosophical and legal questions with almost no operational measurement in 2026 lab practice. Mixing them into a capability roadmap produces heat, not a bottleneck list.

Measurement axes this desk will actually use

If AGI is a spectrum, measurement has to be multi-axis. A single Elo or a single “% of tasks automated” number is a press artifact. The axes below are the ones this series will keep returning to. They are not a new benchmark suite. They are a checklist for reading other people’s suites—and for noticing when a claim has not been measured at all.

Measurement axes for AGI-shaped claims (desk checklist, not a scored index)
Axis Question it asks Typical 2026 proxy (and its failure)
Breadth How many domains without a new stack? Multi-task leaderboards — mix contamination with competence
Transfer / shift Does skill survive a change of surface form? Held-out variants — rarely published with the launch blog
Sample efficiency How much new data to acquire a skill? Fine-tune curves — often omitted; few-shot is not the same
Horizon / compounding How far before the trace is garbage? Agent evals with short tasks — long tasks are expensive to score
Tool reliability Do actions succeed under schema and policy? Unit tests on tools — ignore permission and side-effect cost
Grounding Does the system’s world match the world’s dynamics? Sim success — sim-to-real is the actual axis
Verifiability Can a checker, assay, or proof catch the error? Math/code pass rates — do not export to open-world prose
Cost / watts / latency Is the competence affordable in the loop that needs it? API list prices — hide search, retries, and robot inner loops
Observability Can a human reconstruct what happened? Chat logs — insufficient once tools write state
Governance surface Who can authorize, halt, and audit? Policy PDFs — not the same as runtime enforcement

A claim that does not say which axis moved is not a roadmap update. It is a press release. Builders should steal this table for internal reviews: when a vendor or a lab says “more general,” ask which row, on which harness, at what cost, with what contamination story. Pair that habit with how to read AI leaderboards if you must consume public numbers at all.

A working definition for this series (provisional, revisable)

For Parts 1 and 2, this desk uses a deliberately modest working definition: AGI-shaped competence is the joint ability to acquire, transfer, and execute a wide range of cognitive and (where claimed) physical tasks under distribution shift, over long horizons, with tools or bodies, at a cost that allows the loop to run outside a demo, with traces that can be audited. That sentence is a conjunction. Missing any clause is a gap, not a near-miss. The 2026 present, as Section IV will argue, has impressive disjunctions—systems that are strong on a subset of the clauses—and no public system that satisfies the conjunction in open-world conditions.

This definition is not a standard. It is a filter. It lets us say “this is a scaling result,” “this is an agent-product result,” or “this is a science-verifier result” without promoting any of them to arrival. It also tells Part 2 what to discuss: software that might close horizon and tools; hardware that might close cost and inner-loop latency; scenarios that vary which clause becomes binding first. If a future system meets the conjunction on a public, held-out, multi-axis harness, update the definition last and the evidence first.

II. History: from Dartmouth to the agent present

The AGI conversation is older than the current industry. It is also more repetitive than the current industry likes to admit. Each wave produced a real method, a real overclaim, a real winter or plateau, and a residue that the next wave inherited as its research agenda. Reading that sequence is not antiquarianism. It is how you stop treating 2023–26 as Year Zero. The methods changed. The sociology of hype, funding, and disappointment did not change as much as the hardware did.

Five research waves from Dartmouth and AI winters through expert systems, statistical ML, deep learning, scale, and the 2024-26 reasoning and agents era
History as overlapping waves, not a straight climb: each era leaves a residue the next methods have to attack. The vertical axis is public ambition, not measured generality.

1950s: the proposal that named the field

Two documents still sit under almost every later argument. Alan Turing’s 1950 paper “Computing Machinery and Intelligence” asked whether machines could exhibit intelligent behavior and proposed an imitation-game operationalization that later generations both used and outgrew. The 1956 Dartmouth Summer Research Project on Artificial Intelligence—proposed by John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon—gave the field its name and its original optimism: that “every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it.” That sentence is a research program and a category error waiting to happen. “In principle” is not a timeline. “Precisely described” assumed a kind of formal access to intelligence that later cognitive science and later machine learning both complicated.

Live capture of the Wikipedia page for the Dartmouth workshop, showing the 1956 proposal and organizers McCarthy, Minsky, Rochester, and Shannon
The Dartmouth workshop remains the public origin story. The proposal’s confidence that intelligence could be precisely described is the optimism later winters had to metabolize.

Early work mixed symbolic search, game playing, machine translation sketches, and the first neural-net experiments (including the perceptron line associated with Frank Rosenblatt). The important historical fact is not that these systems were “primitive.” It is that the measurement culture was already brittle: a program that proved theorems in a micro-world or played a restricted game was easy to describe as a step toward general intelligence, because the field had not yet distinguished demo competence from transfer. That distinction is still the live wound in 2026 eval culture. The tools are better. The temptation is identical.

Winters: when the residue became impossible to ignore

“AI winter” is a journalistic phrase for a funding and employment contraction after over-promised capability. The first widely cited contraction in the 1970s followed disappointments in machine translation and in the gap between laboratory micro-worlds and open-ended tasks. The UK Lighthill report (1973) is the usual public marker: a critical assessment of AI research that helped reallocate attention and money. The second contraction, around the late 1980s and early 1990s, followed the boom and bust of commercial expert systems and specialized Lisp hardware. The details differ by country and lab. The mechanism rhymes. A method works in a constrained domain. Institutions extrapolate. The domain boundary does not move as fast as the budget. Trust collapses faster than the method deserved, because the method had been sold as a road rather than as a tool.

Winters are not proof that a method was worthless. Expert systems left behind knowledge-engineering practices, rule engines, and a hard lesson about brittleness. Neural-net winters left behind backpropagation (popularized for multilayer nets in the mid-1980s) and a research community that would later meet GPUs and labeled data. The lesson this desk wants is narrower: when a field confuses a local success with a general theory, the correction arrives as a winter rather than as a paper. A 2026 roadmap that refuses dates is, among other things, an attempt not to write the next winter’s preface.

Expert systems: competence without transfer

The 1970s and 1980s expert-system wave—MYCIN-like medical consultants, XCON-like configuration systems, and a commercial ecosystem around rule bases—demonstrated that encoded specialist knowledge could be operationally valuable. It also demonstrated the residue that symbolic AI never fully escaped: knowledge acquisition was expensive, rules interacted in unanticipated ways, and competence collapsed at the edge of the encoded domain. The systems were, in the language of Section I, deep and narrow. They had verifiability of a sort (you could inspect a rule) and almost no transfer.

Neuro-symbolic work in the 2020s is not a rewind to MYCIN. The modern hybrid usually lets a network propose and a checker verify. But the expert-system residue is still the right warning label: structure you cannot afford to encode, you will not have. Either you learn it from data, or you accept a narrow system, or you find a domain (mathematics, compilers, some sciences) where the structure is already formal. AI-for-science and neuro-symbolic paths in the Frontier taxonomy are, historically, attempts to keep the inspectability of expert systems without paying the acquisition cost in hand-written rules.

Statistical machine learning: prediction without a mind story

From the 1990s through the late 2000s, the center of gravity in many applied fields moved from hand-built knowledge to statistical learning: hidden Markov models and n-grams in speech and language, support vector machines and boosting on tabular and perceptual tasks, probabilistic graphical models as a shared language. The intellectual shift mattered as much as the leaderboards. Intelligence-as-knowledge became intelligence-as-generalization-from-examples. The bitter, useful discovery was that a well-regularized predictor on the right features often beat a brittle theory of the domain.

The residue of this wave is double. First, evaluation became the culture—shared datasets, shared metrics, shared leaderboards. That culture is why 2026 can talk about contamination and saturation at all. Second, the representation was still mostly hand-designed. Statistical ML won by being honest about prediction and limited about features. Deep learning is what happened when the features started to be learned too, at a scale statistical ML’s compute budget could not support. If you only remember one hinge: the field did not jump from Dartmouth to transformers. It spent two decades learning that prediction plus shared evals is a more reliable institution than a theory of mind.

Deep learning, 2012–2019: perception scales, then language follows

The 2012 ImageNet result associated with Krizhevsky, Sutskever, and Hinton (AlexNet) is the public hinge for the modern wave: large labeled data, GPU training, and convolutional networks produced a step-change in visual recognition that the rest of the field could not ignore. The following years stacked architectural and objective-function wins—deeper residual networks, sequence-to-sequence models, attention as a differentiable routing mechanism, and, in 2017, Vaswani et al.’s Transformer, which made parallelism and scale in sequence modeling industrially tractable. BERT-style bidirectional pretraining and GPT-style autoregressive pretraining then showed that language had a self-supervised surplus: text on the internet was a training signal large enough to grow a general approximator.

Game-playing reinforcement learning ran in parallel, not as a side show. DeepMind’s Atari results and the AlphaGo / AlphaZero line (Silver and colleagues) showed that search plus learning plus a well-defined reward could exceed human specialists in closed, perfect-information games. That is an AI-for-science and RL-path ancestor as much as a “deep learning” headline. It also produced a residue that 2026 still lives with: superhuman closed games do not automatically export to messy institutions. The verifier was the rules of the game. Open-world language does not come with that verifier.

By 2019 the field had a new common sense that was also a new overclaim. The common sense: learned representations plus scale plus the right objective move perception and language together. The overclaim: that the remaining gap to general intelligence was mostly more of the same. The 2020–23 wave tested that overclaim at industrial budget.

Scale, 2020–2023: the engine becomes an industry

Kaplan et al.’s 2020 scaling-laws paper gave labs a planning language: loss as a smooth function of parameters, data, and compute. Hoffmann et al.’s 2022 Chinchilla work corrected the compute–data balance and made “undertrained large models” a first-class mistake. Brown et al.’s GPT-3 (2020) made in-context learning a public fact: a large autoregressive model could do new tasks from prompts without a gradient step. Ouyang et al.’s InstructGPT work (2022) and the ChatGPT productization made post-training the interface layer the public actually met. Open-weight lines—especially Meta’s LLaMA family and later competitors—turned access mode into a strategic variable rather than an afterthought. The open-weight versus closed frontier split is a 2020s institutional fact, not a 1950s one.

What scale bought is easy to understate if you are tired of the word. It bought multilingual fluency, non-trivial coding, usable translation, draft-quality analysis, and a single interface that could be pointed at many tasks. What scale did not buy is the residue this article exists to name: guaranteed systematic generalization, grounded physics, reliable long-horizon tool use, cheap unlimited search, and a verifier for open-world claims. The 2022–23 assistant boom productized the first list. The 2024–26 reasoning-and-agents boom is an attempt to chip the second list without abandoning the engine. That is why Scaling / LLMs is tagged Contested on /frontier: the live question is no longer “does scale work?” It is “which scale still pays, on which axis, at what inference bill?”

Reasoning and agents, 2024–2026: the time axis flips

The newest wave is not a new architecture so much as a new place to spend compute. Chain-of-thought prompting (Wei and colleagues, and a large following literature) made intermediate tokens a public technique. Process rewards, verifiable-reward reinforcement learning on math and code, and “reasoning” product modes flipped the scaling axis from train-time parameters to test-time search. A model that thinks longer on a hard item is not automatically more general. It is a system that spends FLOPs where a 2020 chatbot spent charm. The desk’s explainer on reasoning models versus chat models is the product-level map of that split.

Agents, in the 2026 product sense, are the control interface of the same wave. A chatbot returns text. An agent plans, calls tools, observes, and stops under a budget. Most shipped systems are still LLM tool loops with traces of uneven quality—not pure RL agents that improve by acting in the world. That distinction matters historically. The field has wanted autonomous agents since the 1950s. What it has in 2026 is a cheap planner (the language model) bolted to expensive, fragile actuators (APIs, browsers, robots). The practical chatbot-to-agent map is how this desk tells builders to climb that ladder without pretending the top rung is AGI.

Embodiment re-entered the industrial conversation in the same window: vision-language-action models, humanoid fundraising, and sim-to-real as a capital story. World-model language (JEPA-style joint-embedding prediction, Dreamer-style latent dynamics, video generators used as roll-outs) re-entered as a critique of next-token planning. Neither path is new as an idea. Both are newly funded because scale made the missing pieces—planning, contact, compact prediction—feel like the binding constraints rather than like philosophical objections. That feeling may be correct. It is still a feeling until transfer evidence shows up on held-out physical and long-horizon tasks.

Lessons the later paths are still paying for

Historical residues that 2026 paths are still working off
Era What it actually delivered Residue inherited by later work
Dartmouth / early symbolic A named field; search and micro-world competence Confusion of demo with theory; weak transfer culture
Winters A correction against over-promise Oscillation between hype and premature dismissal
Expert systems Valuable narrow competence; inspectable rules Knowledge-acquisition cost; brittleness at the domain edge
Statistical ML Prediction as institution; shared evals Hand-designed features; metrics that can be gamed
Deep learning 2012–19 Learned perception and sequence models; closed-game RL Verifier-rich wins that do not export to open world
Scale 2020–23 A general approximator and an assistant industry Diminishing generic pretrain; eval inflation; inference cost
Reasoning / agents 2024–26 Test-time compute; tool loops as a product category Horizon failure; observability debt; search bills

The through-line is not “progress is an illusion.” Progress is real and lumpy. The through-line is that each wave’s residue is the next wave’s agenda. Scale’s residue is planning, grounding, structure, coordination, verification, and watts. That is exactly the Frontier path list. History did not deduce the eight paths from first principles. It deposited them.

III. A roadmap of lab bets: eight paths as a bottleneck stack

The Frontier taxonomy is the site’s way of not pretending there is one road. Climate tags—Contested, Rising, Core, Accelerating, Niche, Active, Breakthrough, Long-term—are editorial descriptors of how the desk currently reads public activity. They are not probabilities, citations, or investment advice. A path is a bet on which constraint is binding. Labs usually run several bets. Product teams should pick the row that matches their failure mode, not the loudest keynote.

Eight Frontier paths stacked as interacting layers: scaling, world models, RL and agents, embodiment, neuro-symbolic structure, multi-agent coordination, AI for science, and neuromorphic hardware
Read the eight paths as a stack of constraints, not as eight magazines. Residual error on one layer is the research agenda of the others.

Think stack, not menu. Scaling supplies a general approximator. World models, when they work, supply an internal simulator. RL turns prediction into policy. Bodies and tools supply action. Neuro-symbolic structure and scientific domains supply constraints next-token training does not invent. Multi-agent setups multiply specialists. Hardware decides which loops fit energy, latency, and cost. The shorter connector article, how the frontier paths connect, is the coupling diagram for this stack. The rest of this section is the deep dive: what each path claims, what it has actually bought by 2026, what it does not automatically buy, and which table a builder should steal.

Frontier path map — bottleneck each route tries to relieve (desk heuristic, 2026)
Path Site climate tag Primary bottleneck it attacks What it does not automatically buy
Scaling / LLMs Contested Broad competence; data/compute ceiling Reliable long-horizon action; cheap inference; grounded physics
World Models Rising Planning, counterfactuals, compact prediction A body, a reward, a verification loop
RL & Agents Core Turning models into policies and tool loops Correct structure if the simulator or reward is wrong
Embodied AI Accelerating Physical grounding; contact-rich data Unlimited digital scale; easy evals
Neuro-Symbolic Niche Compositionality, constraints, checkable steps Web-scale perception by itself
Multi-Agent Active Specialization, debate, parallel work A source of truth; a smaller failure surface
AI for Science Breakthrough Verifiable domains; synthetic curricula Open-world common sense
Neuromorphic / HW Long-term Energy, latency, on-device loops New algorithms; better rewards

Scaling / LLMs: the engine that still sets the ceiling

Scaling / LLMs dominates because it worked across a shocking range of tasks. Kaplan-style scaling laws and later compute–data balance work gave labs a planning language: parameters, tokens, FLOPs, smoother loss. Pretraining still sets a ceiling; post-training and serving decide what users get. That split is the pretrain → post-train → inference stack, and it is the first thing a non-lab team should name when a vendor says “we scaled.” Most application pain lives in the last two layers. Most AGI rhetoric lives in the first.

The path is Contested because the live question is no longer whether scale works. It is which scale still pays. Generic web pretrain shows diminishing returns on saturated probes. Targeted scale—code, math, multimodal pairs, preference data, and test-time compute—still moves specific curves. Calling the whole path “over” is a category error. Calling it “enough by itself” is the opposite error. Public commentary associated with Yann LeCun and François Chollet argues that current architectures overfit statistics and lack biases for genuine reasoning or efficient skill acquisition. Scaling-first lab leadership replies that capabilities keep appearing and that search at test time is a new dimension. Both can be locally true. Scale is still the cheapest way to buy broad competence—and a weak way to buy guaranteed systematic generalization or physics that survives a robot hand.

Access mode is a separate decision from architecture. Closed APIs concentrate eval hygiene, safety policy, and serving research inside a vendor. Open weights concentrate adaptation, inspection, and geopolitical variance in the downstream ecosystem. Neither mode is “more AGI.” The open-weight versus closed frontier explainer is the access-mode map; this path section is the capability-engine map. Do not fuse them. A smaller open model that you can evaluate and fine-tune can beat a larger closed model on your distribution without moving the AGI spectrum an inch.

Scaling / LLMs — what to watch versus what to ignore
Watch Why it matters Ignore or defer
Data mix and contamination stories Saturated probes stop being evidence Parameter counts as a personality
Post-train and reasoning-mode deltas Users meet this layer, not the base loss Train-FLOP rumors without a model card
Inference price at your actual horizon Test-time search is a new bill “Scale is dead / scale is enough” slogans

World models: planning without treating next-token as a mind

World models are internal predictors that support imagination: if I do X, what happens next, and what would have happened if I had done Y? Ha and Schmidhuber’s 2018 world-model agents, Dreamer-style latent dynamics, and LeCun’s Joint Embedding Predictive Architecture (JEPA) program disagree on machinery. They share a diagnosis: a text completer is a poor planner unless it has compressed a simulator into its weights—or unless you are willing to spend a fortune of tokens reconstructing that simulator in prose every time.

Must that simulator be an explicit module? The world-model camp says agents cannot plan or learn efficiently from sparse action without one. The “prediction is sufficient” camp—often associated in public remarks with Andrej Karpathy and, in another register, Richard Sutton’s emphasis on learning from experience—replies that a good enough predictor already contains a world. The distinction, on that view, is bookkeeping. Builders should ask a narrower question that does not require settling the philosophy: is your system using learned latent dynamics, a physics engine, a video generator as a roll-out, or a language model with tools and a long window? Those four objects have different error modes. Long context is a real capability and a cost center, not compact latent planning. Video systems blur “media generator” and “simulator.” A generator that looks plausible can still violate contact, object permanence, or conservation. Plausibility is not dynamics.

World models couple tightly to RL and embodiment and weakly to support-chat UX. If your product failure is retrieval or tone, you do not have a world-model problem. If your product failure is “the system cannot stick to a multi-step plan that depends on how the environment evolves,” you might. The Rising tag means public momentum and architectural debate, not a completed layer. A world model without a reward is a video. A world model without a body or a tool is a daydream. A world model with both, and with a held-out transfer test, is the actual research object.

World-model design choices — objects that are easy to conflate
Object What it optimizes Typical failure
Latent dynamics (Dreamer-like) Compact roll-outs for policy learning Latent drift; reward misspecification
Joint-embedding prediction (JEPA-like) Predict in representation space, not pixels Unclear export to language-tool products
Video / generative roll-out High-bandwidth imagination Looks right, dynamics wrong
Long-context LLM as “memory of the world” Verbalized state over a window Cost; inconsistency; not compact planning
External simulator + learned policy Ground-truth dynamics in a closed world Sim-to-real; coverage of the real mess

RL & agents: the loop that turns prediction into policy

RL & Agents is tagged Core because product teams already live on it, whether they use the words or not. Reinforcement learning from human feedback (and later AI feedback) sits under every major assistant. Process rewards, verifiable-reward RL on math and code, and search-at-inference squeeze extra competence from a base model. Pure RL agents—systems that improve primarily by acting, not only by imitating text—remain the conceptual endgame even when most shipped “agents” are still LLM tool loops with a thin bandit or a heuristic planner on top.

Sutton’s “Bitter Lesson” is scaling’s ally: general methods that leverage computation win. RL is the other half of that sentence: computation needs a loop and a signal. Without a reward, verifier, or preference model, scale produces better autocomplete. With a bad reward, it produces a better reward-hacker. That is not a slogan. It is the operational history of assistant post-training and of every game-playing result that overfit a proxy. The 2026 product agent—plan, tool, observe, stop—is the control interface through which any more general system will touch the world. It is not AGI. Failure modes are already operational: loops that never halt, unsafe writes, unreadable traces, tools that succeed in unit tests and fail under partial observability. See the practical map and climb only when a single observable loop is green.

RL & agents — three layers that get marketed as one word
Layer What it is What “success” means
Preference / instruction RL Post-train a base model to be usable Helpfulness, format, refusals—not new world knowledge
Verifiable-reward RL + search Spend test-time compute where checkers exist Higher pass rates on math/code-like tasks
Closed-loop agents Act on external state under a budget Task completion with traces, not a demo GIF
Open-ended RL (research) Improve by experience in a rich environment Transfer, not high score on the training simulator

Embodied AI: what a body buys—and what it does not

Embodied AI is tagged Accelerating as capital moved into humanoids, dexterous manipulation, and vision-language-action (VLA) stacks. Public names in the industrial conversation include Figure, Physical Intelligence, 1X, and Boston Dynamics, plus a large academic community. The research claim is older than the fundraising: intelligence that never pokes the world lacks causal grip on it. Perception trained on internet images can name a mug. It does not automatically know what happens when the mug is wet, stacked, or handed to a moving person.

The counter-claim is also public and not silly: digital environments offer huge data at near-zero marginal cost. If the goal is software-native generality—code, law-adjacent text, analysis, design—physical grounding may be a detour with a brutal data-collection tax. Desk synthesis: embodiment is necessary for some AGI-shaped competences (mobile manipulation, household mess, contact-rich tool use) and optional for others (theorem-like reasoning, large-scale software, many scientific computation loops). Sim-to-real gaps mean this path will not outrun LLM pretrain on tokens per day. Its contribution is qualitative: grounding that synthetic text cannot fake. VLA models sit at the junction of scaling, world models, and RL. They are evidence that the chat stack is being pointed at motors. They are not proof that humanoids are the AGI form factor. Form factor is a product bet. Grounding is a competence bet. Do not let a humanoid keynote win the definition.

Embodied AI — competences a body can buy versus detours
Competence Why a body (or a rich sim) helps When a body is the wrong first bet
Contact-rich manipulation Dynamics are not in the text distribution Your product never leaves software
Causal grip on everyday physics Intervention beats observation at some margin You needed a better retrieval system
Human-environment interfaces Houses and factories are built for bodies You needed an API integration
Data that text cannot fake Failure is felt, not only described You cannot afford the collection loop

Neuro-symbolic: structure without a 1990s rewind

Neuro-symbolic work is tagged Niche: smaller paper volume, not “unimportant.” The bet is that neural pattern recognition plus symbolic machinery—logic, programs, graphs, formal solvers—buys compositionality and systematic generalization that scale alone has not delivered. The 1980s rewind fear is mostly a branding problem. The modern form is hybrid: a network proposes; a checker, compiler, SMT solver, or geometry engine verifies. DeepMind’s AlphaGeometry line is the public exhibit in mathematics: neural guidance plus symbolic deduction on large synthetic formal data. Tool-using language models that emit code or proof sketches are cousins of the same idea, even when nobody files them under “neuro-symbolic” in a conference track.

Chollet’s argument—that intelligence is skill-acquisition efficiency, and that many benchmarks over-reward memorized skill—lives mostly here, even when it is phrased as a scaling critique. For operators, the practical form is already on the desk and does not require a unified architecture paper: schema-validated tools, type checkers, policy engines, explicit knowledge graphs, structured outputs with validators. You do not need a neuro-symbolic AGI to use constraints. You need them because unconstrained generation is a known failure mode. The Niche tag means you should not wait for this path to “win” before you put a checker in the loop. It also means you should not expect a logic layer to replace web-scale perception. The hybrid is the point.

Neuro-symbolic patterns already in production-shaped systems
Pattern Neural half Symbolic half
Code agents Propose edits and plans Tests, types, linters, compilers
Formal math / geometry Guide search Proof checkers and deduction engines
Tool calling Choose a tool and arguments JSON schema, authz, idempotency keys
Business workflows Draft and classify Rules, ledgers, human approval gates
Knowledge-graph RAG Embed and retrieve Explicit edges and constraints

Multi-agent: coordination with a cost

Multi-agent systems are tagged Active. Two claims travel under one label and should be separated. The AGI-flavored claim is that generality might emerge from societies: specialization, markets, debate, protocols, cultural accumulation. That is a research question with a long academic history in multi-agent systems and in theories of collective intelligence. The product-flavored claim is modest: split planner from worker, or proposer from critic, when one thread cannot hold the permissions, context, or verification you need. The second claim is already an architecture choice. The first claim is not implied by the second. Shipping a supervisor graph does not mean you have instantiated a society that discovers new competences.

Emergent specialization in simulated populations is interesting and mostly not how enterprise software should be designed in 2026. Supervisor graphs in production are documented in orchestration patterns. More agents usually mean more traces, more cost, and more ways to lose a source of truth. Treat multi-agent as a coordination layer on tool loops—not a substitute for a competent single agent. If you cannot observe one agent, you cannot govern a hundred. If your eval does not beat a well-tooled single agent, the extra personas are theatre. Theatre is a known failure mode of this path and of this decade’s slideware.

Multi-agent — when the extra roles pay
Situation Why split roles When not to
Permissions cannot live in one process Least privilege is an architecture You only needed a better system prompt
Review is cheaper than retry Critic vs proposer can catch schema errors The critic shares the same blind spots and tools
Work is naturally parallel Map-reduce over files or tickets The merge step has no owner
You are studying collective behavior The research object is the society You marketed the study as a product AGI

AI for science: the cleanest training signal

AI for Science is tagged Breakthrough because a few systems already changed their fields in public. AlphaFold (Jumper et al., Nature) reset expectations for protein structure prediction. Later lab and academic work in materials, weather, and mathematics—including the AlphaGeometry line (Trinh et al., Nature)—showed the same abstract recipe in other verifier-rich domains. The AGI-relevant point is not “science is solved.” It is that science offers verifiers: energy functions, assays, proof checkers, reanalysis, held-out experimental confirmation. Those verifiers make reinforcement learning and synthetic data less circular than web-text self-play. A model that invents a plausible molecule still has to face a lab. A model that invents a plausible paragraph often does not.

The path exports artifacts and methods—search plus learning, synthetic curricula, hybrid neural-symbolic loops—that later migrate into general agents. Labs that look “science-only” from a consumer-internet vantage often run the most rigorous version of the recipe other teams want for open-world AGI. That is why this desk treats science as a path on the AGI road rather than as a vertical application. Superhuman closed domains can still be clumsy in messy institutions. Do not read a Nature paper as an AGI demo. Read it as evidence that when the reward is real, the stack compounds. The export question—what transfers from protein geometry or olympiad geometry into open-world agency—is still mostly unanswered. Answering it is more valuable than announcing that science AI “is AGI.”

AI for science — what exports to general agents, what does not
Exportable method Why it might travel What probably does not travel
Synthetic data under a checker Unlimited practice with a grade The specific domain’s ontology
Search + learned guidance Test-time compute with a bound A cheap checker for open-world prose
Closed-loop experiment design Agency with a real reward The lab’s physical infrastructure
Uncertainty that faces measurement Calibration is forced Internet-scale chat incentives

Neuromorphic / hardware: the floor under every other path

Neuromorphic / HW is tagged Long-term because it does not ship the next assistant. It tries to change the physics of running one. Spiking networks, analog and in-memory compute, and brain-inspired chips—public programs include Intel’s Loihi line and IBM’s NorthPole work—target large gains in energy per inference, especially on sparse or event-driven workloads. GPU-centric scaling remains the industrial present. That is a time-scale split, not a contradiction. A long-term path can be correct about the 2030s energy floor while being the wrong procurement for a 2026 product team.

World-model roll-outs, test-time search, robot inner loops, and multi-agent debate are all inference-heavy. If capability per watt does not rise, those loops stay in the lab or in high-margin APIs. Hardware is not a rival theory of mind. It decides which theories run in the wild. For almost all product teams in 2026, the live question is still GPUs, batching, quantization, distillation, and self-host versus API. Waiting on a neuromorphic miracle this quarter is how you fail to ship an eval harness. Ignoring watts entirely is how you discover that your “reasoning agent” is a pricing accident. Part 2 will treat hardware and energy as first-class futures. Part 1 only needs the constraint on the table: every other path has an implicit watts assumption.

Hardware questions that bind other paths (2026 industrial present vs long-term bets)
Loop you want 2026 default substrate Why neuromorphic / custom silicon is a later lever
Chat and short tool use GPU/TPU serving, batching, cache Volume is already optimized on incumbent stacks
Long reasoning traces Same accelerators, worse utilization Energy per useful token becomes the product limit
Robot inner loop Edge GPU or offload to cloud Latency and joules per cycle dominate
Always-on sensing Mostly not shipped at brain-like efficiency Event-driven chips are the research bet

How labs actually mix the paths (not a ranking)

Labs do not pick a path as a sports team. They publish a mix, and the mix moves by quarter. Public research hubs are a better primary source than commentary threads. The screenshots below are live captures of how three large labs present their own research surfaces—not evaluations of their internal roadmaps, and not a claim that unpublished work is absent.

Live capture of Hugging Face Papers, a public research feed used here as a lab-watch proxy for frontier paper traffic
Public paper feed (Hugging Face Papers capture). Desk uses transparent feeds for “what is circulating,” not marketing logo walls.
Live capture of Anthropic’s public research page, showing safety- and model-adjacent publications
Anthropic’s public research surface is often read as scaled models plus constitution-shaped post-training and safety work. Path mix is an editorial reading of what is published, not a private org chart.
Live capture of Google DeepMind’s public research page, showing science, agents, and large-model work
Google DeepMind’s public surface has long included games-as-RL, science, and large models. That combination is a path mix: verifiers, search, and scale in one institution.

A builder’s reading, not a scoreboard: DeepMind-shaped public gravity has included science, games-as-RL, and large models; OpenAI-shaped gravity has included scale, RLHF, and agents; Anthropic-shaped gravity has included scaled models plus constitution-shaped post-training and interpretability-adjacent safety; Meta FAIR-shaped gravity has included world models and open weights; robotics-first firms are on a different binding constraint, not “behind.” Details change by quarter. Use lab watch for release hygiene and this article for the overlay. If you ship agents, you are already on the RL & agents path: climb chatbot → tools → multi-agent only when eval says so. Do not adopt a lab’s mascot path as your identity. You will consume a mix whether you like the branding or not.

IV. The 2026 present: a snapshot, not a scoreboard

A present-tense section is where roadmaps usually start lying. They either underplay what shipped—because the author is tired of assistants—or they promote what shipped into AGI because the author is tired of waiting. This desk will do neither. The 2026 present is a set of productized islands in a larger ocean of residues. The islands are real. The ocean is not a rounding error.

2026 capability snapshot showing strong islands in fluent chat, coding, and short-horizon tools versus gaps in long-horizon agency, physical robustness, and cheap verified open-world action
Capability as islands and gaps, not as a single percentage to AGI. Strong product surfaces sit next to research residues that Part 2 has to take as given.

Capability snapshot (qualitative, multi-axis)

The working definition in Section I was a conjunction. The 2026 snapshot is easier to read as a set of clauses that are independently true or false in public systems. We do not attach unpublished scores. We attach the kind of evidence a non-lab reader can actually inspect: product surfaces, model cards, research blogs, and the failure modes that show up when you try to use the systems as if the conjunction were true.

2026 capability snapshot — islands versus residues (desk synthesis, qualitative)
Clause Public 2026 status What “good” looks like in products What still breaks
Breadth (many domains, one stack) Strong island One assistant or API pointed at many desk tasks Silent failure outside the training-adjacent set
Fluency and draft quality Strong island Usable first drafts in many languages Confident wrongness; style over substance
Code and math with checkers Strong-to-partial Tests and compilers catch a large class of errors Design, specification, and long refactors
Short-horizon tool use Partial island Schema-valid calls, retrieval, light browser work Partial observability; permission edge cases
Long-horizon agency Residue Demos and bounded workflows Compounding error; lost goals; spend blowups
Transfer under shift Residue Fine-tunes and RAG on a known distribution Held-out formats; new tools; new workplaces
Physical grounding Residue / early product Constrained cells, teleop, scripted skills Household mess; contact; sim-to-real
Open-world verification Residue Citations and retrieval as a crutch No checker for most prose claims
Affordable inner loops Partial Chat is cheap enough to productize Search, video, robots, multi-agent debate
Auditable traces Partial Chat logs and some tool logs Governance-grade replay of side effects

Two readings of this table are both wrong. The pessimistic reading treats the residue rows as proof that “nothing happened.” The optimistic reading treats the island rows as proof that “only engineering remains.” What happened is narrower and more useful: a general approximator plus post-training plus short tools became an industry, while the conjunction in the working definition remained a research object. That is a historically large shift and not an arrival.

Eval inflation: why the snapshot cannot be a leaderboard

The statistical-ML wave left shared evals as an institution. The scale wave stressed that institution to breaking. Contamination (test items in training data), saturation (everyone scores high, so deltas are noise), construct leakage (the benchmark stops measuring the thing you cared about), and product-eval mismatch (the user task is not the exam) are now first-class problems. Reasoning modes add a new inflation channel: extra test-time compute can raise a score without raising the transfer or horizon clauses. Multi-agent setups add another: you can spend more roles on a benchmark that a single agent would also pass if you gave it the same tools and budget.

This desk’s rule for 2026 is not “ignore all numbers.” It is do not let a number stand in for the conjunction. A math contest score is evidence about verifiable-reward domains. A coding bench is evidence about repositories that look like the bench. A chatbot arena is evidence about preference in short conversations. None of these is an AGI meter. If you need a hygiene guide, use how to read AI leaderboards and keep a private harness for the tasks you actually buy. Public evals are a weather report. They are not a survey of the ocean.

Productization: what actually shipped

By 2026 the industrial present is not “models in papers.” It is assistants in operating systems and IDEs, retrieval-augmented internal tools, coding agents on bounded repos, voice and multimodal interfaces that are good enough for some workflows, and a small set of computer-use and browser-use products that still live closer to demo than to unsupervised labor. The model stack is how those products are built: a pretrained ceiling, a post-trained personality and tool policy, and an inference layer that decides whether the loop is affordable. Open and closed access modes coexist; enterprises route around both. That routing is the open versus closed story, not the AGI story.

Productization changes the roadmap in two ways. First, it creates feedback that research-only eras did not have: millions of traces of what users try to do, what they accept, and what they retry. That is a post-training resource and a safety resource. It is not automatically a world-model resource, because user chat is a thin slice of the world. Second, it creates a political and economic fact: institutions now depend on systems that are islands, and they will be asked to treat those islands as if they were the conjunction. The gap between product marketing and the residue table is where governance pressure will land—not only where research curiosity will land. A roadmap that talks only about papers will miss the present.

Infrastructure: the silent constraint on every island

The 2026 infra picture is GPU-and-TPU-centric training and serving, with a growing inference-economy literature around batching, caching, quantization, speculative decoding, and distillation. Long context and reasoning traces made serving a first-class research area rather than a deployment afterthought. Data-center power, networking, and the supply of accelerators are now binding constraints on how many experiments a lab can run and how many tokens a product can afford to think with. None of that is neuromorphic. All of it is hardware in the ordinary sense. Part 2 will go deeper on software stacks and hardware futures. Part 1 only needs this: the present is already infra-limited in ways the 2010s were data-limited and the 1990s were theory-limited. A path that ignores serving will look promising in a paper and absent in a product.

Builders who are not labs should read infra as a routing question. If your failure is latency or cost, you are not waiting on AGI. You are in the inference layer: smaller models, caches, routing, and fewer tokens of theatre. If your failure is competence on a private distribution, you are in post-training and data, not in a new accelerator generation. If your failure is that the model cannot act reliably, you are in agents and tools. Infra is the floor. It is not the diagnosis for every pain.

Safety and governance (a sketch, not a doctrine)

Part 1 is not a safety textbook. It needs a sketch because the 2026 present already includes governance as a product constraint, not only as a research conference. Labs publish usage policies, model cards, and (unevenly) system cards. Jurisdictions are writing rules that treat general-purpose models as a distinct class. Enterprises add their own authorization layers because the default assistant will draft things that must not be sent and call tools that must not be open. The measurement-axis row labeled “governance surface” is live: who can authorize, halt, and audit is a capability question once systems write state.

Two errors dominate public safety talk, and both break a roadmap. The first is safety as a synonym for delay: treat every residue as a reason to freeze productization, including the islands that are already in production. The second is safety as a synonym for PR: treat a policy PDF as a substitute for runtime enforcement and traces. This desk’s middle position is operational. If you ship tool loops, you need halt conditions, permission boundaries, and replayable logs before you need a theory of superintelligence. If you claim AGI-shaped agency, you inherit a larger governance surface whether you wanted the politics or not. Part 2 can discuss longer-horizon scenarios. Part 1 only insists that governance is already on the 2026 snapshot, sitting next to eval inflation and inference bills, not in a science-fiction appendix.

Gap list (the brief that Part 2 inherits)

The residues in the snapshot table are the input to Part 2. Naming them here keeps Part 1 from pretending the present is finished:

  • Software-stack gaps. Memory that is not a growing paste of tokens; planners that keep a goal under partial observability; tool ecosystems that fail closed; eval harnesses that measure horizon and shift, not only exams; observability that reconstructs side effects. These are agent and systems problems, not only model problems.
  • Hardware and energy gaps. Test-time search, video-as-simulator, multi-agent debate, and robot inner loops share a watts-and-latency problem. Incumbent accelerators will absorb some of it. Architectural hardware bets try to change the exponent. Neither is a 2026 default for most teams.
  • Verification gaps. Science and code show what a checker buys. Open-world action does not have an equivalent checker. Until it does, “more agency” is also more unchecked side effect.
  • Grounding gaps. Text and video pretraining are not contact. Simulators leak. Household and factory generality remains a data and transfer problem.
  • Scenario gaps. Different clauses may close in different orders. A software-first future, a robot-first future, and a verifier-first future are not the same world. Part 2’s job is to keep those forks visible without assigning them dates or winners.

If you take nothing else from the present-tense section, take this: the islands justify the industry; the gaps justify the research agenda; neither justifies a countdown. The shorter connector showed how paths couple. Part 2 will show how software and hardware might close—or fail to close—specific gaps. This Part 1 stops when the gaps are named clearly enough to hand over.

FAQ

What is AGI Roadmap Part 1 actually claiming?

That you can map history, definitions, and eight Frontier paths as a bottleneck stack without announcing a date, a winner, or a single architecture. The claim is methodological: residues, not prophecies, should drive the agenda.

How is this different from the shorter “how frontier paths connect” article?

The connector is the coupling diagram you can read in one sitting. Part 1 is the long brief: spectra and exclusions, measurement axes, a historical wave from Dartmouth to 2026 agents, and a present-tense snapshot. Use the connector for routing; use Part 1 when you need the archive and the definitions.

Will Part 2 predict when AGI arrives?

No. Part 2 discusses software stacks, hardware and energy constraints, and scenario families. It inherits this article’s refusal to publish a countdown. If a future edition adds dates, it will have stopped being this series.

Is scaling still the main path in 2026?

It is still the main engine for broad competence. It is contested as a complete theory of general intelligence. Track diminishing generic-pretrain returns separately from post-training, data quality, and test-time compute. See where returns diminish.

Do I need world models in my product this year?

Only if planning or dynamics is the bottleneck—not retrieval, tone, or schema-valid tool use. Most software products should ship tools, evals, and a competent base model first. A video generator is not automatically a planner.

Are multi-agent systems a shortcut to AGI?

There is no public evidence that extra agents substitute for better models, rewards, or verifiers. They can help specialization and review. They raise observability cost. If one agent is unobservable, a society is ungovernable.

Does AGI require a robot body?

Required for what? For household-mobile manipulation, a body or a very good stand-in is part of the competence. For software and many scientific loops, the case is weaker. Do not let one definition win by slogan. Embodiment is a path, not a membership test.

Which labs are winning the AGI road?

This desk does not rank a winner. Labs publish different path mixes. Watch access modes, evals, what shipped, and which bottleneck a paper actually touched. Use lab watch 2026 for release hygiene, not for a trophy table.

Why not treat a high reasoning-bench score as AGI?

Because the working definition is a conjunction. A score can move verifiable-domain competence via test-time compute without moving transfer, horizon, grounding, cost, or auditability. Exam-style probes are also exposed to contamination and saturation.

What should a non-lab team do with this map?

Do not pick a path as identity. Consume APIs and open weights from labs that already mix scaling, RL, and tools. Put leverage in post-training on your data, retrieval, evaluation, and agent scaffolding. Competing on pretrain scale is a capital-structure decision, not a branding decision.

Where do safety and regulation sit on this roadmap?

On the 2026 present, as a live constraint on tool loops and as a growing institutional layer around general-purpose models. They are not a substitute for measurement, and they are not delayed until a fictional arrival date. Runtime halt, permissions, and traces are the builder-facing piece.

How should I read climate tags on /frontier?

As editorial descriptors of public activity and discourse—Contested, Rising, Core, and the rest—not as probabilities or investment grades. When the public mix changes, the tag should change after the evidence, not after a keynote.

What to do next

If you need the one-page stack picture, read how the Frontier paths connect. If you allocate research or capital, stay on this page long enough to steal the measurement-axis table and the residue list, then go to Part 2: software, hardware, and futures. That article takes the gaps named here—software loops, watts, verification, grounding, and scenario forks—and treats them as design problems rather than as a calendar.

If you ship product this quarter, do not wait for Part 2 to become a checklist. Start with the chatbot-to-agent map for the control interface, the 2026 model stack for where your pain actually lives, and reasoning versus chat before you pay for test-time theatre. Return to the Frontier path map when a new paper claims to have ended the race—and check which bottleneck it actually touched. If you want to argue with the map in public, the community is the place to post a held-out eval, not a date.

Maps beat countdowns. Subscribe for Frontier Desk notes on labs, agents, and scaling—framed as bottlenecks, not prophecy. Part 2 continues with software, hardware, and futures.

Subscribe to the Everything is AI newsletter · Frontier paths · Community

Sources

  1. Turing, Computing Machinery and Intelligence (Mind, 1950) — operational imitation-game framing; later generations both used and outgrew it.
  2. Dartmouth workshop (1956) — public origin account — McCarthy, Minsky, Rochester, Shannon proposal that named the field.
  3. Lighthill, Artificial Intelligence: A General Survey (1973) — public marker of the first widely cited UK contraction.
  4. Krizhevsky, Sutskever, Hinton, ImageNet classification with deep convolutional neural networks — 2012 hinge for the modern perception wave.
  5. Vaswani et al., Attention Is All You Need — Transformer as the industrial sequence-model substrate.
  6. Kaplan et al., Scaling Laws for Neural Language Models — classic pretrain scaling framing.
  7. Hoffmann et al., Training Compute-Optimal Large Language Models (Chinchilla) — compute–data balance.
  8. Brown et al., Language Models are Few-Shot Learners (GPT-3) — in-context learning as a public scale result.
  9. Ouyang et al., Training language models to follow instructions with human feedback — RLHF as public assistant post-training.
  10. Sutton, The Bitter Lesson — computation-leveraging methods versus hand-built structure.
  11. Ha & Schmidhuber, World Models — compact latent models for policy learning.
  12. LeCun, A Path Towards Autonomous Machine Intelligence — JEPA-oriented world-model program (OpenReview position paper).
  13. Silver et al., Mastering the game of Go with deep neural networks and tree search (Nature) — search plus learning in a verifier-rich closed game.
  14. Jumper et al., Highly accurate protein structure prediction with AlphaFold (Nature) — AI-for-science verification loops.
  15. Trinh et al., Solving olympiad geometry without human demonstrations (Nature) — neural guidance plus symbolic deduction.
  16. Wei et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — intermediate tokens as a public technique.
  17. IBM Research on NorthPole — public on-chip inference efficiency program.
  18. Intel neuromorphic computing (Loihi) — public spiking / neuromorphic research line.
  19. OpenAI Research — public research surface (live capture in-article).
  20. Anthropic Research — public research surface (live capture in-article).
  21. Google DeepMind Research — public research surface (live capture in-article).

What we did not test: We did not train models, run a new AGI-style battery, rank labs, score unpublished systems, or treat rumor names as facts. Site climate tags are editorial descriptors of the Frontier taxonomy, not probabilities. Historical sketches rely on standard public accounts; they are not an archival monograph. Capability language in Section IV is qualitative desk synthesis, not a private leaderboard.

Corrections: When a lab’s public path mix changes, when a verifier-heavy or embodied stack produces documented transfer evidence, or when a measurement axis gets a better public harness, update the coupling, the snapshot table, and the as-of date first—not the title’s metaphor and not a countdown. If a source URL moves, replace the link; do not invent a replacement paper.