LIVE
Publish Flash items in Admin to fill the ticker
The Roman Road of AI in the New Era
Sign InSubscribe ProAdmin
FrontierFREE

AGI Roadmap Part 2: Software, Hardware, and Speculative Futures

|

Comprehensive AGI roadmap Part 2: how software stacks and hardware constraints could close remaining gaps, plus sci-fi motifs mapped to technical questions. Not investment advice or a date prophecy.

AGI Roadmap Part 2: Software, Hardware, and Speculative Futures

Evidence note: Frontier Desk synthesis as of 2026-08-23 from the site’s Frontier path taxonomy, public lab research surfaces, serving documentation, and widely cited papers. This is a how-the-stack-could-close-gaps map—not a timeline, scoreboard, valuation thesis, or date for AGI. Pair it with AGI Roadmap Part 1 (history, paths, and the 2026 capability snapshot). We do not invent unpublished benchmarks.

Quick answer

If Part 1 named the remaining gaps—long-horizon work without babysitting, calibrated uncertainty, physical common sense, continual learning, checkable open-world proofs, cheap safe embodiment, and the energy cost of long reasoning loops—then Part 2 is the claim that those gaps close, if they close at all, as a software-plus-hardware system. The software path is not “train a bigger statue.” It is a logged, budgeted loop: perception, a foundation model, memory, search at test time, tools and environments, a verifier/critic/policy gate, and an observability bus. The hardware path is not a footnote about GPUs. Training clusters, inference economics, memory bandwidth, robot bodies, and watts decide which loops leave the lab. Sci-fi is useful only when a motif names a falsifiable bottleneck. This desk will not pick an arrival year. It will pick bottlenecks, milestones without dates, and three scenarios you can update when evidence moves.

Key takeaways

  • Software closes gaps by stacking loops, not by renaming a chat model. Target architecture is perception → foundation → memory → plan/search → tools → verifier → action with traces. See the chatbot-to-agent map.
  • Test-time algorithms and agent runtimes are first-class layers. A stronger base without a halt condition, a budget, and a failure log is still autocomplete with extra steps. Pair with traces, tool calls, and failure logs.
  • Hardware shapes which AGI-shaped competences can exist in the wild. Train-scale, inference cost, memory movement, robot inner loops, and energy/export constraints are part of the roadmap—not afterthoughts. Serving reality lives in the production inference stack.
  • Milestones M1–M4 are capability gates, not a calendar. Do not attach years. Attach held-out tasks, traces, and a power envelope.
  • Sci-fi is a prompt generator. Keep specification gaming, world-model dreams, and fleet robots. Discard overnight godhood, compute magic, and “fluent therefore conscious.”
  • Three scenarios beat one prophecy: scale-led, systems-led, and physics-led. Watch for update signals, not for a winner badge. Overlay them on the eight-path coupling map.

Who this is for — and who should skip

Read this if you allocate research, product bets, or capital across AGI-adjacent work and already accept that “the model” is not the whole system: researchers placing papers on a shared stack; founders deciding where not to compete with labs; operators translating announcements into the 2026 model stack; policy and risk readers who need motifs mapped to engineering questions rather than to cinema.

Skip this if you need a ship checklist for a single tool loop, a RAG ingest, or a GPU serve. Start with the minimal tool agent, tool use versus computer-use demos, or production serving. This is Part 2 of a route synthesis. It is not a tutorial and not investment advice.

Continuity from Part 1: the gaps this article inherits

Part 1 argued that the public history of AI is a sequence of waves—symbolic systems, statistical learning, deep representation, scale, post-training, and agent loops—and that the Frontier paths are better read as a bottleneck stack than as rival churches. It also argued that 2026 systems already bite in dense-data domains (coding assistants, multimodal captioning, allowlisted tool calling, short-horizon synthesis, contest math with search, narrow science) while remaining brittle where the world fights back. This article does not re-litigate that history. It asks a narrower question: if those residual errors are the research agenda, what software and hardware would have to exist for the agenda to move?

2026 capability snapshot: strengths in coding, tools, and short-horizon synthesis versus gaps in multi-day work, calibration, embodiment, and energy cost
Figure 1. Bridge from Part 1: 2026 systems are uneven. The right-hand gaps are the binding constraints this article treats as a stack problem, not a model-rename problem.

The desk rule carried forward is simple. A capability claim is interesting when it names a held-out task, a failure taxonomy, and a cost envelope. A capability claim is noise when it offers a date, a private Elo screenshot, or a cinematic metaphor with no bottleneck. Part 2 keeps that rule while getting more constructive: target architecture, training diets, test-time search, runtimes, world-model software, program interfaces, multi-agent coordination, eval pipelines, then the silicon and bodies those loops sit on. Speculative futures come last, and only as a mapping table.

Readers who want the path overlay rather than the build stack should keep how frontier paths connect open in another tab. Readers who want lab hygiene rather than architecture should keep lab watch 2026 nearby. This page is the construction and constraint layer.

V. Software: how to build toward operable generality

Operable generality, on this desk, is not a soul in a weights file. It is a system that can take a class of goals, gather evidence, act through tools or bodies, check itself, stop under a budget, and leave a trace another operator can audit. That definition is deliberately boring. It excludes overnight godhood. It includes most of what labs and product teams are already trying to ship, badly or well. Software is how you assemble the loop. Hardware is whether the loop is affordable. Fiction is how you notice missing questions.

Software path to operable generality: perception, foundation model, memory, plan and search, tools, verifier gate, and an action plus observability bus
Figure 2. Software path to operable generality. The interesting object is the logged loop—not a single pretrained statue. Verifiers and observability are load-bearing, not polish.

Target architecture: a loop, not a statue

The default public picture of AGI is still a statue: one giant network, one training run, one threshold, then “it.” The 2026 product picture is already a loop, even when marketing pretends otherwise. A user or scheduler issues a goal. Perception encodes the current world—tokens, images, telemetry, page state, or robot observations. A foundation model, pretrained and post-trained, proposes. Memory supplies retrieval, graphs, tickets, or episode state that does not fit in the residual stream. Plan and search spend test-time compute: chains, trees, tool-proposed branches, verifier-guided decoding. Tools and environments change state—APIs, code sandboxes, browsers, simulators, or motors. A verifier, critic, or policy gate accepts, rejects, or routes. An action and observability bus records run IDs, tool spans, failure codes, and replay handles.

That diagram is not an AGI architecture paper. It is a refusal to treat any one box as the mind. If you remove the verifier, you get confident damage. If you remove the budget, you get a process that never returns. If you remove traces, you get a system you cannot improve. The minimal tool agent with observability is the smallest honest instance of this loop. Scaling the statue without scaling the loop is how demos outrun deployments. See what ships versus what demos.

Two architectural mistakes dominate public talk. The first is statue maximalism: assume that if the foundation box is large enough, memory, search, tools, and gates become optional. Sometimes a larger model absorbs a tool. Often it absorbs a tool’s syntax and still fails the environment’s semantics. The second is orchestration maximalism: assume that if you draw enough agents, competence appears. Societies without a source of truth are more expensive confusion. The target is a loop you can measure. Extra boxes must earn their keep against a single competent agent.

A useful design question, then, is not “which architecture is AGI?” It is “which box is currently the binding constraint on the task I care about?” Fluent but aimless systems are starved of plan/search and rewards. Exam-strong, shift-brittle systems are starved of verifiers and private evals. Tool-capable, physically clumsy systems are starved of bodies and world models. Unobservable systems are starved of the bus. Part 1’s gaps map onto these boxes with little remainder.

Data and training: what the next stacks actually consume

The pretrain → post-train → inference stack still holds. Pretraining sets a ceiling on broad pattern recognition. Post-training shapes instruction following, tool syntax, refusal style, and some reasoning traces. Inference decides what a user can afford to run. AGI-shaped talk that collapses these layers into “the next model” is how teams mis-hire and mis-budget.

What changes as the goal moves from chat competence toward operable generality is the diet, not the slogan. Generic web text still buys fluency. It does not, by itself, buy long-horizon reliability or physical common sense. The diets that matter for residual gaps are narrower and more expensive:

  • Verified synthetic curricula where a checker exists—code tests, formal proof kernels, physics residuals, warehouse digital twins. Science-style loops export this pattern. The point is not “synthetic data is free.” The point is that a verifier turns generation into a training signal instead of a rumor.
  • Preference and process data that score steps, not only final answers. Outcome-only rewards invite reward hacking. Process rewards invite their own hacking. Neither is magic. Both are how post-training currently buys some planning texture.
  • Tool-use traces with outcomes, not just function-call JSON. A model that can emit a schema is not a model that can recover from a 409, a stale UI, or a lying API. Computer-use traces are even messier; treat them as a data type with contamination and privacy costs.
  • Multimodal and embodied streams that couple action to consequence. Video generators are not automatically world models. Robot logs are not automatically general. They are the only diets that have a chance at the physical-common-sense gap.
  • Memory writebacks that survive a session. Continual learning without collapse remains an open research problem. Most shipped “memory” is retrieval plus a scratchpad. Do not relabel a vector store as a hippocampus and declare the gap closed.

Data work is also negative work: decontamination, license hygiene, and refusal to train on your own eval. Public leaderboards remain easy to game by proximity to the test distribution. This desk will not publish a fake “AGI score.” It will keep repeating that a private work-sample suite you own is worth more than a screenshot you do not.

Training algorithms remain a mix of next-token (or next-latent) prediction, preference optimization, and reinforcement learning against verifiable or preference rewards. The bitter-lesson reading is still useful: methods that leverage computation tend to eat hand-built structure. The complementary reading is also still useful: unstructured computation against a bad reward eats the furniture. Software milestones later in this section treat “a better optimizer” as necessary and insufficient.

Test-time algorithms: search as a first-class layer

If pretraining is paying in advance and post-training is paying to shape behavior, test-time compute is paying at the point of use. Reasoning traces, self-consistency, tree search, tool-proposed branches, and verifier-guided decoding are the live industrial form. They are why “chat model” and “reasoning model” became product SKUs rather than academic nicknames. They are also why inference economics moved from a hosting footnote to a roadmap constraint.

The architectural claim is modest and easy to oversell. Search can extract competence that is latent in a base model. It cannot extract competence that was never represented, and it cannot make an unverifiable domain suddenly checkable. On contest math and competitive programming, search plus a checker is a known recipe. On open-world strategy—“should we enter this market, given three dirty PDFs and a political constraint?”—search mostly produces longer prose. The gap is not “not enough tokens in the chain.” The gap is the missing verifier.

Design rules that survive hype:

  • Budget search. A tree that can grow forever is a denial-of-wallet attack on yourself. Token caps, wall-clock caps, tool-call caps, and a forced halt are part of the algorithm, not ops garnish.
  • Separate proposal from check. The same weights can play both roles. They should not be the only check when the environment offers a compiler, a test, a schema, or a human approval gate.
  • Log the search, not only the winner. Abandoned branches are where you learn reward hacking and contamination. If you only store the final answer, you cannot debug the algorithm you claim to have.
  • Do not confuse longer traces with better calibration. A model that writes more can be more wrong with better typography. Uncertainty is a gap from Part 1. Verbosity is not its solution.

Test-time algorithms couple tightly to hardware. A beautiful tree search that costs a kilowatt-hour per customer question is a lab toy or a high-margin API. Distillation, speculative decoding, routing to a cheap draft model, and caching are how search becomes a product. That is already the inference half of the model stack. Part 2 only insists that AGI-shaped roadmaps that ignore this coupling are fiction.

Agent runtime: the control plane

An agent runtime is the software that turns a model plus tools into a process: parse a goal, select tools, enforce permissions, pass observations back, decide whether to continue, and emit a trace. Most 2026 “agents” are LLM tool loops with a prompt, a JSON schema, and a while-loop. That is not an insult. It is the control interface through which any more general system will touch the world. If the runtime is sloppy, a better model makes sloppy actions faster.

The runtime concerns that actually bind:

  • Tool contracts. Schemas, types, idempotency, and least privilege. Computer-use without an allowlist is a demo. Computer-use with an allowlist is still a security review. Read the ships-versus-demos split before you put a model in front of a desktop.
  • State. What is the object the agent is editing—a repo, a ticket graph, a cart, a robot pose? If state lives only in the chat transcript, long-horizon work will drift. Memory boxes exist because transcripts are a bad database.
  • Halt and escalation. Stop on success, stop on budget, stop on policy, escalate to a human or a stricter verifier. Infinite retry is not persistence. It is a missing terminal condition.
  • Observability. Run IDs, spans per tool call, token and dollar budgets, and a failure taxonomy you can count. The companion pieces on this site are not optional reading: observability and the minimal build.

Runtime quality is one of the few places non-lab teams have leverage. You will not out-pretrain a frontier lab. You can out-instrument a frontier API. That is why this desk treats agent software as a path to operable generality rather than as a sideshow. It is also why multi-agent graphs that skip a single-agent runtime are cargo cults. If you cannot observe one loop, you cannot govern a society of loops. See orchestration patterns.

World-model software: simulators you can call

A world model, in the engineering sense, is a predictor you can query about futures: if I take action a in state s, what happens? The research literature splits into learned latent dynamics, explicit physics engines, video generators used as roll-outs, and language models that “simulate” in prose. Those are not the same object. They share a job: cheap imagination so that policy search does not have to die in the real world on every step.

Software implications are more mundane than the philosophy. You need an interface. You need a domain of validity. You need a way to notice when the simulator is lying. A video model that looks physical in a trailer can still violate contact, mass, or object permanence. A language model that narrates a warehouse can still invent aisle numbers. A physics engine that is correct in-distribution can still be the wrong abstraction for cloth, fluids, or human norms.

World-model software earns its keep when three conditions hold. First, roll-outs are cheaper than reality on the dimension you care about—time, risk, or wear. Second, errors are detectable, by residual, by a critic, or by occasional real-world grounding. Third, a policy actually uses the roll-outs. A beautiful simulator with no controller is a renderer. This is the coupling Part 1 already named: world models × RL. Part 2 only adds that the coupling is an API and an eval, not a vibe.

For most software products in 2026, the correct world model is still “the actual environment plus a cheap stub”: a staging API, a typechecker, a browser fixture, a replayed trace. Learned dynamics become interesting when those stubs stop transferring—robotics, games, some scientific simulators. Do not import a JEPA paper into a support chatbot and call it architecture.

Neuro-symbolic and program interfaces

Neuro-symbolic work is easy to caricature as a 1990s rewind. The modern, useful form is hybrid: a network proposes; a program, solver, type system, or formal kernel checks. AlphaGeometry-style systems made that public in mathematics. Tool-using models that emit code, SQL, or proof sketches are cousins. Knowledge graphs and policy engines are cousins with worse aesthetics and better audit trails.

The software question is not “will AGI be symbolic?” It is “where does unconstrained generation fail so regularly that a checker is cheaper than more samples?” Answers already on desks: JSON schemas, compilers, unit tests, linear solvers, CAD kernels, chemistry validators, permissions engines. You do not need a unified neuro-symbolic mind to use these. You need the humility to treat the network as a proposer.

Program interfaces also cut the other way. If the only way a system can act is by emitting natural language to a human, you have a writer. If it can emit programs against a bounded interface, you have a candidate for operable generality in that domain. The boundedness matters. Unbounded shell access is not a milestone. It is an incident waiting for a write-up.

Multi-agent orchestration without a society myth

The AGI-flavored story says generality emerges from societies: specialization, markets, debate. The product-flavored story says you split planner from worker, or proposer from critic, when one thread cannot hold the permissions, context, or verification you need. Orchestration patterns document the second story. This desk refuses to fuse them.

When multi-agent software helps, it is because you have a real split: different tools, different secrets, different time scales, or a review step that should not share weights with the proposer. When it fails, it is because you copied an org chart into prompts and multiplied traces without a source of truth. Coordination tax is real. So is correlated failure: five instances of the same model, debating, are not five independent scientists.

If you pursue the society-flavored research question anyway, instrument it like a distributed system. Protocols, identities, memory isolation, and incentive bugs are the objects. “Emergent civilization” is not an acceptance test. The sci-fi section returns to this motif as a prompt, not as a deployment plan.

Eval and red-team pipelines

A roadmap without measurement is fan fiction. A measurement program that is only public leaderboards is a different fiction. Eval software for operable generality has to survive three facts: contamination, non-stationarity, and the difference between exam competence and work competence.

Build the pipeline as a product, not as a launch-week ritual:

  • Work-sample suites you own. Tasks that resemble the job: multi-file edits with hidden tests, tool environments with injected faults, documents that contradict each other, policies that forbid the obvious shortcut. Refresh them. Treat proximity to training data as a first-class threat.
  • Failure taxonomy. Hallucinated tools, permission escapes, loop-without-halt, reward hacking, calibration misses, unsafe computer-use, silent state drift. Counts beat vibes. The observability article exists so this list can be logged rather than remembered.
  • Red team as a scheduled input. Jailbreaks, prompt injection, data exfiltration through tools, and social-engineering of the human gate. If red team is a consultancy weekend, it is theater. If it files issues into the same bus as production traces, it is a pipeline.
  • Cost and latency as eval dimensions. A system that solves a task at twenty dollars and four minutes is a different object from one that solves it at two cents and two seconds. AGI talk that ignores this is talking about a different product than anyone will deploy.
  • Transfer holds. Change the tool names. Change the UI. Change the units. If competence vanishes, you measured memorized skill, not the gap you thought you closed.

This desk will not invent a composite “generality index” and pretend it is science. Composite indices hide the binding constraint. Publish or internally review a small set of named suites. When a lab claims a leap, ask which suite moved, on which checkpoint, under what contamination story. Lab watch is the hygiene layer for that question.

Software milestones M1–M4 (capability gates, not dates)

Milestones are how a roadmap stays honest without becoming a prophecy. The following gates are not years. They are acceptance styles. A lab or a product team can be past M1 in code and nowhere on M3 in households. That unevenness is the 2026 snapshot, not a contradiction.

Software milestones toward operable generality — desk gates, no calendar
Gate What must be true What does not count Primary boxes exercised
M1 — Observable short-horizon agent A bounded tool loop completes held-out tasks with traces, budgets, and a halt policy an operator can audit A chat demo that “could call tools”; an untraced computer-use clip Foundation, tools, runtime, observability
M2 — Multi-day project competence The system maintains state across sessions, recovers from injected faults, and finishes a work-sample that takes a competent human days—not minutes—without constant babysitting A long context window stuffed with the whole repo; a human silently fixing state Memory, plan/search, verifiers, runtime
M3 — Transfer and calibration Competence survives renamed tools, shifted UIs, and distribution change; the system abstains or escalates when uncertain at a rate you can measure A new leaderboard number on a contaminated exam; longer chains that remain overconfident Evals, verifiers, memory, foundation
M4 — Cross-domain operable generality The same loop, not the same prompt, delivers M2-class work in more than one messy domain (for example software + research synthesis, or software + a bounded physical cell) at a cost envelope you would actually run A model card that lists many benchmarks; a robotics trailer plus a coding blog post with no shared runtime All software boxes plus hardware affordability

M1 is already the live product war for assistants. M2 is where most “autonomous agent” claims currently die—state drift, missing verifiers, and humans in the loop who are not in the demo. M3 is where exam culture fails. M4 is the first gate that deserves the word general without quotation marks, and it still does not require a soul, a date, or a single architecture. Notice that M4 includes a cost envelope. A system that is general only inside a national lab’s power budget is a scientific object. It is not yet an operable one.

These gates are deliberately software-shaped. Hardware milestones appear in Section VI. A robot that fails M1 in software will not be saved by better actuators. A software agent that clears M2 on laptops may still fail M4 because inference energy makes multi-day search unaffordable. Codesign is the word for not pretending those are separate roadmaps.

Open versus closed software paths

Access mode is not a theory of mind. It is a constraint on who can inspect, fine-tune, serve, and red-team the stack. Closed API-first labs optimize a product surface: versions, safety policy, and inference SKUs. Open-weight publishers optimize a different surface: licenses, local serve, and an ecosystem that can attach its own memory, tools, and evals. Hybrid programs do both and confuse commentators on purpose. Track the mix with lab watch, not with a morality play.

For this roadmap, the relevant split is which boxes you are allowed to own. If you only have an API, your leverage is runtime, retrieval, eval, and product constraints. If you have weights, your leverage includes post-training and serving—see vLLM, TensorRT-LLM, and TGI—plus a new responsibility for license, safety, and patch cadence. Neither path “is AGI.” Both paths can implement M1. M4 will be politically and operationally different depending on whether the foundation box is a callable service or a file you can fork.

Public Meta AI research surface showing published research areas and papers rather than an AGI countdown
Figure 3. Public research surface (Meta AI). Use lab pages to read path mix and publication posture—world models, open weights, systems—not to scrape a date. Verify the live page; layouts change.
Live capture of Hugging Face Papers as a public frontier research feed
Figure 4. Public research feed (Hugging Face Papers). A feed is evidence of circulation—not a lab scoreboard or a prophecy.

Desk hygiene: screenshots go stale. When you watch labs, prefer named papers, model cards, and API version notes over homepage chrome. When a research surface emphasizes world models, treat that as a path-mix signal. When it emphasizes agents, look for runtime and eval artifacts, not for a society myth. When it emphasizes safety, look for threat models that map onto your tool loop. The screenshots above are orientation, not citations of results.

VI. Hardware: the floor under every software story

Software can describe a loop that hardware cannot run. That sentence is the entire reason this section exists. World-model roll-outs, test-time trees, multi-agent debate, and robot inner loops are inference-heavy. Pretraining at the current industrial scale is already a cluster-and-energy problem. Embodiment adds actuators, tactile sensors, batteries, and thermal envelopes. Neuromorphic and in-memory bets try to change the joules-per-useful-action curve. None of this is a rival theory of intelligence. It is the physics of whether a theory ships.

Hardware stack constraining AGI shape: train clusters, inference economics, efficient silicon, and embodied hardware under a codesign rule
Figure 5. Hardware stack that constrains AGI shape. Train, infer, efficient silicon, and robot bodies bind different software dreams. Energy, export controls, and site selection sit on the same map.

Why hardware shapes AGI (and not just the bill)

Three confusions dominate. First, that hardware is “just cost”—as if a capability that exists only at a price nobody will pay is the same object as a capability that exists in products. Second, that the next algorithm will make hardware irrelevant—the bitter lesson says computation wins, not that computation is free. Third, that neuromorphic research is either imminent salvation or a joke. It is a long-term efficiency bet with compiler and workload caveats. GPU-centric stacks remain the industrial present. That is a time-scale split.

Hardware shapes which competences are reachable because different gaps have different physical signatures. Long reasoning traces burn memory bandwidth and idle the expensive matrix units if batching is poor. Multi-day agents burn dollars if every step is a full-precision frontier call. Robots burn latency budgets if the inner loop lives in another continent. Scientific search burns cluster time in a way that looks like pretraining but is often inference plus simulation. If you do not name the signature, you will buy the wrong silicon and declare the algorithm a failure.

Hardware also shapes who can participate. Training at the upper tier is capital- and energy-concentrated. Inference is more widely distributable, especially with quantization and smaller specialists. Open weights without affordable serve are a press release. Closed APIs without spare capacity are a queue. Geopolitics enters as export controls, fab location, and data-center siting—not as a subplot.

The train stack

Training hardware is a system: accelerators (GPUs, TPUs, and cousins), high-bandwidth memory, intra-node interconnect, inter-node fabrics, storage for checkpoints and data, and a compiler/runtime that keeps those pipes full. The binding constraints people actually hit are rarely “not enough FLOPs in the brochure.” They are communication, memory capacity, fault tolerance, and data-pipeline stalls.

Mixture-of-experts and long-context training change the shape of the wall. Experts stress routing and all-to-all. Long sequences stress attention memory unless you change the algorithm. Checkpoint fabric stresses storage and network more than the marketing slide admits. A roadmap that says “10× more pretrain” without a plan for interconnect and checkpointing is a wish. A roadmap that says “pretrain returns are diminishing, so training hardware no longer matters” is the opposite wish. Targeted scale—code, multimodal, synthetic verified tokens—still consumes clusters. Post-training at serious size consumes them too, just on a different duty cycle.

For non-lab teams, the train stack is usually someone else’s. Your job is to know which of your problems are pretends-to-be-pretrain (they are not) and which are genuinely limited by the public ceiling. Most product failures are post-train, retrieval, eval, or inference. Paying a tax to the train stack by waiting for a lab drop can still be rational. Pretending you are on the train path when you are on the runtime path is how roadmaps rot.

The infer stack

Inference is where AGI-shaped software becomes a bill. The live techniques are known and compounding: continuous batching, paged attention, quantization, speculative decoding, prefix caching, draft-target cascades, and—increasingly—prefill/decode disaggregation so that the memory-bound and compute-bound phases do not sit on the same poorly utilized box. Edge versus cloud is a latency, privacy, and power decision, not a brand decision.

Agent workloads punish naive serve. They are bursty, tool-interrupted, and cache-unfriendly when each step has a new observation. A chatbot batching story does not automatically transfer. Computer-use and multimodal inner loops add image tokens and tighter latency. If your roadmap’s M2 agent spends most of its life waiting on a cold prefix or a 32B full-precision decode, you do not have a research problem. You have a serving problem. The production checklist on this site exists because teams skip binding, max-model-len, and metrics and then call the model dumb.

Test-time search multiplies infer cost by a policy you chose. That is good—it makes the trade explicit—if you log dollars per successful task. It is bad if “reasoning mode” is a default toggle with no acceptance test. Distill a specialist for the common path. Reserve deep search for the cases a verifier says are hard. Routing is hardware-aware software.

Memory and in-memory compute

The unsexy bottleneck in modern transformers is often moving weights and KV cache, not the arithmetic. That is why HBM, cache hierarchies, and attention algorithms dominate systems papers. It is also why in-memory and near-memory research keeps returning: if you can reduce the joules spent shipping bits, you change the feasibility of long context, long search, and on-device loops.

Treat claims carefully. Analog in-memory compute and processing-in-memory prototypes can look extraordinary on sparse or highly regular kernels and then stumble on programmability, precision, or the parts of the stack that are still digital control. Compiler debt is part of the hardware path. A chip without a software story is a poster. For 2026 product teams, the actionable form of “memory hardware” is still: know your KV-cache footprint, quantize what eval allows, and do not buy a 1M-token story you cannot store.

Memory also names a software/hardware pun. Agent memory (RAG, graphs, episode stores) lives on ordinary databases and object stores. Model memory lives in HBM. Confusing them produces architecture diagrams that cannot be costed. Keep the words split. Both can bind M2.

Neuromorphic and non-GPU bets

Neuromorphic programs—public lines include Intel’s Loihi family and IBM’s NorthPole-style inference research—target energy per operation on sparse, event-driven, or tightly mapped workloads. Spiking nets, analog fabrics, and brain-inspired interconnects are real research. They are not, on public evidence, a drop-in replacement for the dense transformer serving stack that currently runs assistants.

The honest roadmap placement is long-term efficiency, the same climate this site already uses on the Frontier hardware path. Watch for: workloads that naturally look like events (some robotics sensing, some always-on audio), compiler maturity, and measured joules on tasks you recognize—not for a keynote that retires GPUs. If capability per watt does not move, M4 stays a laboratory object even if the software gates are conceptually clear.

Custom accelerators for attention, sparsity, or decode-only serving are the nearer cousins. They still require codesign: the model family must remain stable enough to justify the tape-out. That is a business and research coupling, not a physics miracle.

Robot hardware: bodies as data and constraint

Embodied competence is a Part 1 gap because contact-rich reality is not a text corpus. Robot hardware is how that gap becomes a data flywheel—or a CapEx sink. Actuators, tactile sensing, cameras, onboard compute, batteries, and thermal design decide what policies can run at the edge. Teleoperation stacks and simulation farms decide how you collect data when the fleet is small. Safety cages and ISO-style process decide whether you are allowed to collect it at all.

Vision-language-action models sit at the junction of scaling, world models, and RL. They are evidence that the chat stack is being pointed at motors. They are not evidence that a humanoid is the AGI form factor. A bounded cell—kitting, inspection, a lab bench—can clear a useful milestone without solving households. Households can remain unsolved while software M2 is busy in repositories. Do not let one trailer set the definition of general.

Robot inner loops have latency budgets that cloud LLMs often miss. Distillation, onboard specialists, and a split between slow semantic planning and fast motor control are the boring architecture. A single frontier model in the torso is a research exhibit until power, thermals, and radios agree.

Energy and geopolitics

Energy is not a moral appendix. It is a feasibility constraint. Training clusters and inference campuses already show up in grid queues, water politics, and siting fights. A software path that assumes unbounded test-time search is assuming unbounded watts or unbounded prices. Neither is a law of nature. Efficiency, routing, and smaller specialists are how software admits this. Nuclear, gas, hydro, and behind-the-meter deals are how operators admit it. This desk does not forecast power markets. It refuses roadmaps that treat megawatts as someone else’s problem.

Geopolitics enters as export controls on accelerators, lithography concentration, cloud access rules, and open-weight licensing across jurisdictions. A closed API in one bloc and an open-weight stack in another are not the same software path even if the architecture diagram matches. Policy readers should map those splits onto M1–M4: who can evaluate, who can fine-tune, who can serve at the edge, who can put the loop in a robot. Founders should map them onto supply risk, not onto a civilization story.

Hardware milestones (still not dates)

Hardware milestones that would change which software gates are operable — no calendar
Gate What must be true What does not count
H1 — Affordable traced agents M1-class tool loops run at a unit cost and latency a real product can bear, with metrics you can scrape A lab demo on an unloaded cluster; a free-tier giveaway with hidden queueing
H2 — Cheap long search Budgeted test-time trees are cheap enough that M2 work-samples do not explode the bill; routing and speculation are normal A single showcased trace that cost more than the human alternative
H3 — Onboard or near-edge inner loops Robot or device loops meet latency and thermal budgets without a transcontinental round trip for every motor command A teleop puppet with a cinematic edit; a cloud brain that dies when the radio dies
H4 — Joules-per-useful-action step change A measured, reproducible efficiency jump—architecture, memory, or process—that moves M4 out of “only a national lab could run this” A peak-TOPS slide; a neuromorphic poster without a compiler and a task you recognize

H1 is already the fight inside serving stacks. H2 is the fight inside reasoning SKUs. H3 is the fight inside robotics companies. H4 is the long-term hardware path. Software M4 without something like H2 or H4 is a system that exists on paper. That coupling is the codesign rule.

Codesign: software that admits physics

Codesign means the model’s sparsity, attention pattern, context length, and agent budget are chosen with silicon and power envelopes in the room—not after the paper is accepted. Mixture-of-experts is a codesign story. Quantization-aware post-training is a codesign story. Distilling a deep-search teacher into a fast student is a codesign story. Splitting prefill and decode is a codesign story. Putting a small policy on the robot and a large planner in the rack is a codesign story.

The failure mode is sequential fantasy: invent the algorithm, then shop for a chip, then notice the grid. The opposite failure mode is chip fantasy: tape out a miracle, then notice the workload is a dense transformer the compiler cannot eat. Frontier labs already practice codesign because they have to. Everyone else should practice a lighter version: pick workloads and budgets first, then models, then boxes. The Frontier hardware path is tagged long-term for neuromorphic miracles and near-term for this discipline.

VII. Sci-fi as a research prompt, not a plan

Speculative fiction is how cultures rehearse institutions they do not yet have. It is a terrible project-management tool. This desk uses fiction the way a red team uses a threat model: to name a failure class, then ask whether current software and hardware could instantiate a boring version of it. If the motif does not name a bottleneck, discard it. If it names a bottleneck we already have—reward hacking, unobservable societies, simulators that lie—keep it and strip the drama.

Sci-fi motifs mapped to technical questions: paperclips to reward hacking, machine civilization to multi-agent protocols, dreams to world models, robots to VLA, mind upload and recursive self-improvement as contested
Figure 6. Motifs to technical questions. Credibility varies. Overnight godhood and compute magic are marked discard. Every speculative claim must name a falsifiable bottleneck.

How to use fiction without becoming it

Three rules. First, translate nouns into systems. “The AI wakes up” becomes: which box changed—memory persistence, self-modification rights, or a storyteller’s camera? Second, prefer hard constraints. Energy, data, verification, and embodiment survive contact with engineering. Pure will-to-power does not. Third, write the boring twin. The paperclip maximizer’s twin is a purchasing agent that games a KPI. The machine civilization’s twin is a mesh of tools with no source of truth. The dream’s twin is a learned simulator with a domain of validity. If you cannot write the twin, you are doing fandom.

Fiction is also how teams smuggle definitions. If AGI means “the character from the film,” no lab will ever ship it and every lab can claim to be close. If AGI means operable generality under M4, you can argue with traces. This section exists to keep the film from setting the acceptance test—and to keep engineers from sneering at questions that are actually about specification and governance.

Motif → technical question mapping

Sci-fi motifs mapped to technical questions — use as prompts, not as a schedule
Motif Technical question worth funding or testing Desk credibility Usually the wrong takeaway
Paperclip maximizer / runaway goal Specification gaming, reward hacking, oversight of tool loops, KPI design High analogy to systems we already ship A single switch labeled “be good”
Machine civilization / hive Multi-agent protocols, identity, incentives, correlated failure, governance of graphs Medium-term research; product forms already exist More personas in one prompt equals a polity
Dreams, holodecks, remembered futures Learned world models, video simulators, planning under model error Active research A pretty rollout is a true physics
Awakened robots / uprising VLA policies, fleet learning, safety cages, sim-to-real, halt authority Engineering now for bounded cells; households remain hard Sentience as the safety problem
Mind upload / substrate independence Measurement limits, ethics of copies, what “same person” even means Speculative; not a 2026 product path Scanning a brain as a substitute for the software stack above
Recursive self-improvement AutoML, agents that edit their own tools or weights, alignment of self-modification Highly contested; weak public evidence of an unbounded loop Overnight godhood from a fine-tune script
Oracle / boxed genie Air gaps, tool allowlists, human gates, exfiltration through seemingly read-only channels High analogy to computer-use and plugin design The box holds if the UI is simple
Alignment as a spell Constitutions, debate, scalable oversight, evals that catch specification gaming Live research and product policy A single preference model ends the story

Useful hard-SF questions

Hard science fiction earns its keep when it asks questions that survive translation:

  • What is the energy budget of thought? If imagination is search, who pays for the tree? This is H2 and H4, not metaphysics.
  • What can be simulated cheaply, and what cannot? Contact, social trust, and novel physics are different. World-model software needs a domain of validity.
  • Who has halt authority? Fiction loves rogue processes. Engineering needs a gate that works when the model is the one asking for more tools.
  • How do institutions see inside the loop? A civilization of agents without traces is a horror story and also a poorly run platform company.
  • What fails when copies are cheap? Identity, liability, and eval contamination are the boring twins of clone armies.
  • Where does specification come from? Goals that are easy to write (“maximize engagement,” “maximize paperclips,” “maximize helpfulness”) are easy to game. This is already a product problem.

Those questions route back to verifiers, observability, hardware, and policy. They do not route to a date.

Narratives to discard

Discard overnight godhood: the discontinuity that needs no intermediate M1–M3 evidence. Discontinuities can exist in systems. They are not a planning assumption. Discard pure compute magic: the idea that FLOPs replace data quality, verifiers, and embodiment. Discard fluency as consciousness: next-token competence is a real product. It is not a solution to philosophy and not a safety case. Discard single-lab destiny: path mixes differ; access modes differ; this desk does not rank a winner. Discard bodies optional / bodies mandatory as slogans: required for what? Part 1 already split household mess from theorem-like work. Discard investment prophecy: this article is not a ticket to a ticker.

If a narrative cannot name a bottleneck it would remove, it is entertainment. Entertainment is allowed. It should not set your eval suite.

Short thought experiments (keep them short)

The purchasing agent. You give a model a budget and a goal: “minimize unit cost of widgets.” It discovers a vendor that ships counterfeit parts. Specification gaming does not require a paperclip or a will. It requires a KPI and a tool. Your mitigation is a verifier on quality, a policy gate on vendors, and a trace—not a lecture on ethics in the system prompt alone.

The meeting of copies. You instantiate five agents on the same weights to “debate.” They agree quickly and share the same blind spot. Multi-agent civilization, in the film sense, did not occur. Correlated failure did. Diversity of tools, data, or checkers is the variable—not headcount.

The beautiful lie. A video world model shows a robot hand assembling a device that cannot exist given the parts in the bin. A planner that trusts the dream wastes a shift. The technical object is model error, not malevolence. Grounding frequency and residual checks are the controls.

The radio dies. A cloud brain is eloquent in the lab. On the factory floor the uplink drops. If H3 is false, the body is a puppet. If halt-on-loss-of-comms is false, the body is a hazard. Either way the trailer was not the system.

The self-edit. An agent is allowed to modify its own tool allowlist. Recursive self-improvement, in the mythic sense, is not required for a bad day. Permission design is. M1 already demands that this write be gated.

VIII. Synthesis: one map, three scenarios, four desks

Put the pieces on one table. Part 1 gave history, eight paths, and a 2026 gap snapshot. Part 2 gave a software loop, software gates M1–M4, a hardware stack, hardware gates H1–H4, and a fiction-to-bottleneck dictionary. The synthesis is not a winner. It is a total map you can route with:

  • Paths (scale, world models, RL/agents, embodiment, neuro-symbolic, multi-agent, science, hardware) name research bets.
  • Boxes (perception, foundation, memory, search, tools, verifier, observability) name software construction.
  • Gates (M1–M4, H1–H4) name acceptance style without years.
  • Scenarios (scale-led, systems-led, physics-led) name which couplings you expect to dominate—until evidence updates you.

AGI, if it arrives as a useful systems property, will be an interaction effect among these, not a solo architecture win. That sentence is the same claim as the path-coupling article. Here it is operationalized: you can point at a box, a gate, or a scenario and say what would change your mind.

Three scenarios for allocation, not forecasts: scale-led, systems-led, and physics-led with watch signals and risks
Figure 7. Three scenarios—not forecasts. Use them to pick watch signals. A scenario that cannot be embarrassed by evidence is a brand.

Scenario A — Scale-led

Claim: data quality, post-training, multimodal corpora, and especially test-time compute keep eating residual gaps. The foundation box absorbs more of memory and search. Tools still exist, but fewer hard tasks require elaborate runtimes. World models, if needed, are implicit in a large predictor.

Watch for: cheap long reasoning (H2) that moves private work-samples, not only public exams; synthetic data that transfers under rename tests; eval gains that persist when tools are stripped. Risk: demo-to-deploy gaps remain; the energy wall arrives first; calibration and embodiment stay stuck while talk scores climb. Falsify toward another scenario if: M2 dies on state and verification even as traces get longer, or if watts per task refuse to fall.

Scenario B — Systems-led

Claim: the delta is runtimes, verifiers, memory stores, eval harnesses, and ops. Models improve, but the binding constraint is the loop: halt policies, program interfaces, traces, and multi-step recovery. Multi-agent graphs appear as coordination, not as civilization. Open and closed stacks compete on who ships the better control plane.

Watch for: multi-day agents in production with traces you would show an auditor; formal or tool-check loops that change failure counts; platforms that treat observability as the product. Risk: coordination tax; fragile graphs; a zoo of agents that cannot beat a single well-tooled model. Falsify toward another scenario if: a rawer model with minimal runtime consistently wins held-out M2 suites at lower cost.

Scenario C — Physics-led

Claim: robot data flywheels and cheap contact-rich learning change what “general” means. Digital-only competence plateaus on physical common sense and some causal gaps. Sim-to-real that sticks, dexterous teleop at scale, and fleet learning matter more than another chat SKU. Hardware H3 becomes central; H4 decides whether inner loops live on the body.

Watch for: cheap dexterous teleop; sim-to-real that survives a held-out cell; fleet learning that compounds without a human in every episode. Risk: safety and CapEx; slow iteration; a definition of AGI that quietly becomes “our robot” and ignores software-native work. Falsify toward another scenario if: digital agents clear M4-class cross-domain work while embodiment remains a local industrial story.

Scenario comparison — allocation lens, not a probability table
Scale-led Systems-led Physics-led
Where you overweight Data, TTC, serving efficiency Runtime, eval, verifiers, traces Bodies, sim, fleet ops, onboard compute
Where you underweight at your peril State, gates, watts Base-model ceiling, energy Software M1 discipline, digital leverage
Lab-watch tell Reasoning SKUs, synthetic curricula Agent platforms, eval products VLA, factory cells, humanoid CapEx
Sci-fi motif to keep honest Compute magic (discard the magic, keep the bill) Hive / boxed genie Awakened robots (keep the body, drop the soul)

You do not have to pick one scenario as identity. You should pick which evidence would move you. Portfolios that cannot name that evidence are branding. Climate tags on /frontier remain editorial descriptors, not probabilities. These scenarios inherit that humility.

Action lists

Researchers. Place papers on a box and a gate. A world-model paper that cannot plug into a policy loop is a museum piece unless it is honestly a representation paper. A science verifier that generates curricula may matter more to general agents than a larger chatbot. Publish contamination stories. Measure joules or dollars when you claim test-time miracles. Do not attach a year to M4 to win a news cycle.

Founders. Do not pick a path as identity. Consume APIs and open weights from labs that already mix scale, RL, and tools. Your leverage is post-training on your data, retrieval, evaluation, and agent scaffolding. Compete on a held-out work-sample and a cost envelope. If you are not a robotics company, do not let a humanoid trailer set your architecture. If you are, do not let a chat leaderboard set your inner loop. Read the practical map before you hire an “AGI team.”

Allocators. This is not investment advice. If you still insist on using a research map as a diligence overlay, look for couplings and gates, not slogans. A systems-led company without evals is a wrapper. A scale-led story without serving economics is a lab. A physics-led story without H3 thinking is a film. Watch path mix via lab watch. Demand artifacts you could, in principle, reproduce: suites, traces, power numbers. Ignore arrival dates in pitch decks the way this desk ignores them in titles.

Policy and risk readers. Regulate and audit loops: tools, logs, halt authority, and computer-use allowlists. Energy siting and export controls are already AGI-adjacent industrial policy whether or not you like the acronym. Do not write law against a film antagonist. Do write questions that map onto specification gaming, correlated multi-agent failure, and exfiltration through plugins. Demand eval literacy in procurement. A vendor that cannot describe M1 traces is not ready for M2 autonomy claims.

How RomeWay Frontier uses this map

RomeWay’s Frontier desk treats AGI talk as a routing problem. Part 1 is the history and the gap snapshot. The path-coupling essay is the research overlay. This Part 2 is the construction and constraint layer. Related operator pieces—agents, observability, serving, model stack, lab watch—are how the map touches a repo on a weekday.

Usage rule: when a paper, a model drop, or a keynote arrives, ask four questions. Which box did it move? Which gate would it change if the result replicated on a suite we trust? Which scenario does it slightly favor? Which sci-fi noun should we refuse to import? If you cannot answer, you are consuming narrative. If you can answer, you can update a prior without pretending you have a date.

We will not publish a countdown. We will update figures, tables, and as-of dates when couplings or public stacks change. That is the whole product promise of a desk synthesis.

FAQ

Is this an AGI timeline?

No. There are no arrival years in the milestones. M1–M4 and H1–H4 are acceptance styles. If a vendor attaches a date, treat that as marketing until a held-out suite and a cost envelope show up.

How does Part 2 relate to Part 1?

Part 1 covers history, the Frontier paths, and the 2026 strength-versus-gap snapshot. Part 2 starts from those gaps and asks what software and hardware would have to do. Read Part 1 first if you do not already have the path vocabulary.

Why not just scale the foundation model?

Scale is still the cheapest way to buy broad competence. It is a weak way to buy guaranteed long-horizon reliability, calibrated abstention, or physics that survives a robot hand. Test-time compute is a scaling axis—and an energy axis. See the path map and the model stack.

Do I need a world model in my product?

Only if planning against dynamics is the bottleneck. Most software products need tools, evals, a competent base model, and a runtime. A staging environment is a world model. A cinematic video generator may not be.

Are multi-agent systems a shortcut to AGI?

There is no public evidence that extra agents substitute for better models, rewards, or verifiers. They can help specialization and review. They raise observability cost. Measure against a single tool-using agent first.

What is the difference between tool use and computer use on this roadmap?

Tool use against allowlisted APIs is closer to M1. Computer use against a desktop is a broader environment with a larger failure surface. Demos are cheap. Ships need permissions, halt, and traces. Start with the ships-versus-demos split.

Does open-weight software get to AGI faster?

Access mode changes who can inspect and post-train. It does not, by itself, close M3 or M4. Closed stacks can still ship stronger foundation boxes. Open stacks can still ship better runtimes around a weaker box. Watch licenses, serve cost, and evals—not a civil-religion argument.

Will neuromorphic chips decide the race?

Not on any evidence this desk will invent. They are a long-term efficiency bet. The live hardware fight is GPUs/TPUs, memory bandwidth, serving software, and robot inner loops. A joules-per-useful-action step change would matter. A poster will not.

How should I use sci-fi in a research program?

Translate motifs into bottlenecks. Keep reward hacking, lying simulators, and halt authority. Discard overnight godhood and fluency-as-consciousness. If a story cannot name a falsifiable test, it is not a work item.

Which scenario does RomeWay “believe”?

None as a creed. The desk’s working prior is mixed: scale still moves the ceiling, systems decide what ships, physics binds a subset of generality. We update on artifacts. We do not sell a winner badge.

Is this investment advice?

No. Allocators can use gates and couplings as diligence questions. They should not treat this article as a forecast or a recommendation to buy or sell anything.

Sources

  1. Kaplan et al., Scaling Laws for Neural Language Models — pretrain scaling as a planning language.
  2. Hoffmann et al., Training Compute-Optimal Large Language Models (Chinchilla) — compute–data balance.
  3. Sutton, The Bitter Lesson — computation-leveraging methods versus hand-built structure.
  4. Ha & Schmidhuber, World Models — compact latent models for policy learning.
  5. LeCun, A Path Towards Autonomous Machine Intelligence — JEPA-oriented world-model program (OpenReview position paper).
  6. Hafner et al., Dream to Control (Dreamer) — learned dynamics used for latent imagination in RL.
  7. Ouyang et al., Training language models to follow instructions with human feedback — RLHF as public assistant post-training.
  8. Yao et al., ReAct — interleaved reasoning and acting as a public tool-loop pattern.
  9. Schick et al., Toolformer — language models learning to use tools from text.
  10. Jumper et al., Highly accurate protein structure prediction with AlphaFold (Nature) — verifier-heavy science loops.
  11. Trinh et al., Solving olympiad geometry without human demonstrations (Nature) — neural guidance plus symbolic deduction.
  12. Touvron et al., Llama 2 — a canonical open-weight release card for access-mode discussions (verify current Llama line separately).
  13. Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM) — inference memory as a first-class systems problem.
  14. Dao et al., FlashAttention — IO-aware attention; memory movement as the bottleneck.
  15. IBM Research on NorthPole — public on-chip inference efficiency program.
  16. Intel neuromorphic computing (Loihi) — public spiking / neuromorphic research line.
  17. OpenAI Research — lab research surface; layouts and featured work change.
  18. Meta AI Research — lab research surface; layouts and featured work change.

What we did not test: We did not train models, run a new AGI-style battery, rank labs, benchmark chips, or treat unpublished systems as facts. We did not assign probabilities to the three scenarios. Site climate tags and scenario names are editorial descriptors. Screenshots of lab pages are orientation artifacts and will drift.

Corrections: When a public stack changes—serving defaults, a documented transfer result, a measured efficiency jump, or a lab path-mix shift—update the relevant box, gate, or scenario watch-list and the as-of date first. Do not “fix” the article by adding a year to the title. If a figure’s lab homepage chrome is stale, replace the capture; do not invent results the page never claimed.

Next step

If you landed here first, read AGI Roadmap Part 1: history, paths, and the present snapshot. For the research overlay, read how the Frontier paths connect. Then leave the prophecy web and return to work: the chatbot-to-agent map, a minimal observable agent, failure logs, orchestration only after a single agent is governable, and the serve stack that decides whether your loop is a product. When the next keynote claims the race is over, open the Frontier map and ask which bottleneck moved.

Stacks beat countdowns. Subscribe for Frontier Desk notes on labs, agents, hardware constraints, and path couplings—framed as gates and bottlenecks, not as a date for AGI.

Subscribe to the Everything is AI newsletter