H
Howardism
Howardism · Vol. 03Plate II · No. 02

Superintelligence Trajectory, in order.

Notes20DomainSuperintelligence TrajectoryOpen Qs49Newest16 Jul 2026Oldest7 Jun 2026

Recursive self-improvement, scaling limits, and the path to ASI.

Map of Content for the superintelligence-trajectory domain — 20 concepts. The path from AGI to ASI: recursive self-improvement, intelligence-explosion dynamics, ASI theory and limits, and frontier governance. Curated entry point; see Home for all domains.

  • The Abstraction Barrier — Lerchner's hypothesis that AI trained on human concepts may be unable to discover genuinely novel conceptual primitives from raw data — capping single instances near AGI — and the embodied bottleneck that grounds concept validation in real-world experiment speed, converting recursive self-improvement into a process paced by empirical science
  • Advantages of Digital Intelligence — The six properties (Table 1) that follow from knowing an AI's source code — I/O speed, processing speed, working memory, substrate independence, lossless replication, high-bandwidth experience sharing — each of which scales with compute in ways biological intelligence cannot, widening the human–AI gap
  • AGI-to-ASI Pathways — DeepMind's four non-exclusive, parallel technological routes from human-level AGI to superintelligence — scaling, algorithmic paradigm shifts, recursive self-improvement, and multi-agent group agency — plus the six frictions (data wall, economics, paradigm-insufficiency, research-gets-harder, abstraction barrier, deliberate slowdown) whose impact is the report's central set of open research questions
  • AI Accelerating AI Development — The empirical core of When AI builds itself: measured evidence AI already speeds AI R&D at Anthropic — >80% of merged code Claude-authored, ~8× code/engineer/day vs 2024, a kernel-optimization eval going 3×→52× in a year, an automated researcher recovering 97% of a weak-to-strong gap, and model next-step judgment beating humans 64%
  • AI R&D Autonomy Evaluation (AECI) — How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives recursive self-improvement; tracked via the AECI capability index plus concrete shortcomings vs. human researchers; Opus 4.8 sits below the frontier and is not close to substituting for research staff
  • Artificial Superintelligence (ASI) (hub) — DeepMind's informal characterization of ASI as a system that exceeds large, well-coordinated human-expert collectives across virtually all domains — distinct from human-level AGI below it and the incomputable Universal AI limit above it, all points on the Legg–Hutter intelligence continuum
  • Autonomous Scientific Discovery — Mythos-class models now conduct novel science with limited human input — autonomous protein/drug design (~10× faster, matching skilled humans), molecular-biology hypotheses preferred ~80% over Opus-class (one E. coli mechanism independently corroborated), and week-long genomics that beat a Science-published model at 100× smaller; the wet-lab analogue of AI-driven formal proof search, and fresh evidence in the research-taste debate
  • Capability-Gated Model Fallback — Fable 5's safeguard architecture: classifiers detect cyber / bio-chem / distillation queries and route the response to a less-capable model (Opus 4.8) instead of refusing — 'fallback, not refusal'; >95% of sessions never trigger; conservative tuning, robust to 1,000+ hours of jailbreak testing; a new point on the safeguard spectrum for capabilities past a risk threshold
  • Effective Compute Scaling — DeepMind's framing of compute growth as ~10×/year of 'effective compute' — the product of hardware improvement (~1.5×/yr), compute investment (~2.5×/yr), and algorithmic efficiency (~3–6×/yr) — and the data-wall and economic frictions that determine how long the scaling pathway to ASI can be sustained
  • Frontier Pause Verification — The arms-control problem of a credible, verifiable slowdown or pause of frontier AI: detectability is harder than for other technologies (training runs are easier to conceal than missile silos), so the Anthropic Institute aims to build the verification systems a multilateral pause would require
  • Fundamental Limits of ASI — Even far-superhuman AI is bound by hard physical (Landauer, Bremermann, Bekenstein, light-speed), complexity-theoretic (P vs NP), and logical (Gödel, Halting) limits — but these negative results are often 'vacuous' in practice because good heuristic approximations exist below the worst case
  • Intelligence Explosion Dynamics — The growth-curve question behind recursive self-improvement: whether AI-accelerating-AI produces exponential, super-exponential/hyperbolic (singularity-in-finite-time), or S-curve dynamics — and the four mechanisms (genetic, cultural, cooperative, data) plus the physical/economic frictions that bound it
  • Multi-Agent Collective Intelligence — DeepMind's fourth pathway to ASI: superintelligence as an emergent property of many coordinated AGI agents — group agents, virtual agent economies, and centrally-steered super-collectives — governed by hoped-for 'multi-agent scaling laws' and the open question of when a homogeneous LLM collective actually becomes more than the sum of its parts
  • Open-Weight Elicitation Irreversibility — A wiki-drawn synthesis of Brown and Gemma 4: if dangerous capability scales with inference budget, then an open-weight release fixes the model's safety evaluation at one budget forever while leaving elicitation budget unbounded and recall impossible — the closed-weight mitigations (classifier fallback, suspension, retention) all require a server the vendor controls
  • Recursive Self-Improvement (hub) — An AI system autonomously designing and developing its own successor; Anthropic Institute's When AI builds itself argues AI is already accelerating AI development (engineers ship ~8× more code/quarter) and lays out three futures — stalled-but-diffused, compounding-efficiency, and full RSI
  • Research Taste as the Human Bottleneck — The narrowing human role as AI absorbs execution: choosing which problems matter, which results to trust, and when an approach is a dead end; the top rung of the autonomy ladder, and the open question of whether taste is 'just another capability' AI fails at then masters
  • Researcher Uplift from Code Output — Thomas Kwa (METR) translates Anthropic's reported 8× code-per-engineer-per-day into serial researcher uplift with production functions: Cobb-Douglas gives U = M^β = √8 ≈ 2.83, CES stays within ±3% of that across elasticities because 8 ≈ e², and a low-stakes-code-discounted model still lands [2.33, 2.66] — so researcher uplift from coding agents alone is plausibly >2×, reconciled with Anthropic's 'well short of 2× overall R&D uplift' because R&D speedup also depends on compute (Greenblatt: labor^0.55 × compute^0.45)
  • Responsible Scaling Policy Evaluations — Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misalignment; the Opus 4.8 determination is that it does not advance the frontier beyond Mythos Preview and that catastrophic risk remains low given current mitigations
  • Transformative Creativity — Boden's three-level model of creativity (combinational, exploratory, transformative) used to locate today's AI achievements — Move 37, AlphaFold, theorem-proving — at the exploratory level within human-given conceptual spaces, and to frame Boden level-3 (creating new conceptual spaces, à la Hassabis's 'could AI rediscover general relativity?' test) as a hallmark requirement of true ASI
  • Universal AI (AIXI) (hub) — Hutter & Legg's formal upper bound on machine intelligence: AIXI, the incomputable agent optimal on average over all computable environments under Solomonoff's universal prior; the theoretical endpoint of the intelligence continuum that ASIs approximate from below

Open questions 49 open

  • Advantages of Digital Intelligence
    • Does training on human data suffice to give digital intelligence human-grade abstractions, or does the low embodiment factor cap concept formation? (The crux shared with [[abstraction-barrier]].)
  • AGI-to-ASI Pathways
    • For each friction: is it a *fundamental blocker* (multi-year plateau) or a mere *friction* (slows, doesn't halt)? The report's central unresolved question. **Synthesized with Anthropic:** [[wiki/derived/rsi-growth-curves-which-friction-binds]] — data-wall and research-gets-harder demote themselves into compute; economics and neural-paradigm are pathway-conditional; the abstraction barrier is the candidate fundamental (re-pacing) blocker; and deliberate slowdown is the only *exogenous* friction — the one Anthropic wants to install and this report doubts can be made to bind.
    • Do the four pathways compound multiplicatively when run in parallel, and how would we detect that early?
  • AI Accelerating AI Development
    • The W2S result didn't transfer to production-scale models. Is that a temporary scaling artifact or a structural limit on autonomous research?
    • The next-step judgment trend (51%→64%) is measured only on weak-human-move slices. What does the curve look like on a representative sample of research decisions?
  • AI R&D Autonomy Evaluation (AECI)
    • "Not close to substituting for senior researchers" is a subjective, internally-sourced judgment. What objective signal would replace it as models approach the threshold?
    • AECI is a single scalar fork of an external index; how sensitive is the 155.5 / frontier-not-advanced conclusion to the choice of the n=11 evaluation set?
    • The shift to "direct measurement of AI R&D acceleration and researcher uplift" is announced but not yet operationalized in this card — what does that measurement look like? **Sharpened:** [[researcher-uplift-from-code-output]] — one external answer: translate a measured code-output multiplier into serial researcher uplift with a production function (Cobb-Douglas/CES), preferring code *output* over per-hour *uplift* because output prices in time reallocation. It also splits the target quantity in two — *serial researcher uplift* (labor only) vs Anthropic's *overall R&D speedup* (labor × compute) — so a rigorous internal measure must state which it reports.
  • Artificial Superintelligence (ASI)
    • Can we even *recognize* ASI? We lack benchmarks for general superhuman performance (only narrow ones like chess), and the tasks must be abstract/open-ended enough to reveal it.
    • Is the jaggedness of capabilities a fundamental theoretical property, or an artifact of comparing against human performance? (Open question 6d in the report.)
    • Where does practical ASI plateau relative to the hard limits — how much slack is there?
  • Autonomous Scientific Discovery
    • Science's verification gap: the formal-proof loop self-validates; here a wrong-but-confident hypothesis costs a wet-lab cycle to falsify. Does autonomy without a fast verifier *increase* the verification bottleneck rather than relieve it?
    • If hypothesis-generation is genuinely at ~80% preference, how much of "research taste" is left as a distinctively human function — and how would you measure the residue?
  • Capability-Gated Model Fallback
    • The >95%/<5% figures are session-level; what's the false-positive rate for *legitimate* security researchers and biologists, whose benign queries are exactly the ones most likely to trip the conservative classifiers?
    • Fallback-not-refusal preserves UX but means the *real* general-access model for security/bio-adjacent work is Opus 4.8, not Fable — does that quietly cap Fable's value for whole professional segments until the trusted-access programs open?
    • The UK AISI's "progress toward a universal jailbreak" is disclosed but not quantified — and the post-launch **access suspension** (see [[claude-fable-5]]) raises the question of whether a safeguard failure forced it.
    • Does swapping to a weaker model on flagged topics create an exploitable oracle (probe which queries trigger fallback to map the classifier's boundary)?
  • Effective Compute Scaling
    • When does more compute reliably yield more *intelligence* — only for some problem classes, or generally? Can quantitative and qualitative scaling be traded off?
    • Can data generation (synthetic, simulated, interactive) actually keep pace with model-size growth, or does the data wall bind first?
  • Frontier Pause Verification
    • What does an AI-training "verification regime" concretely consist of — compute-accounting, datacenter inspection, hardware attestation, on-chip telemetry? The essay names the problem, not the mechanism.
    • Detectability < verifiability: can detection even be made reliable when training runs leave no physical signature and inputs are dual-use?
  • Fundamental Limits of ASI
    • Can we develop theory for "hard *and* inapproximable" problem classes — the only negatives with practical bite?
    • How much slack sits between these fundamental limits and the *practical* ceiling of AGI/ASI systems?
  • Intelligence Explosion Dynamics
    • Can "recursive improvement scaling laws" be formulated — predicting self-improvement curves (and their plateau point) from early-onset datapoints?
    • How far can a *fixed* model's performance be pushed with test-time search alone, and under what conditions does recursive distillation degenerate vs. compound?
    • Which binds first — algorithmic ceilings, the embodied bottleneck, or compute/energy supply — determining exponential vs. hyperbolic vs. S-curve? **Synthesized:** [[wiki/derived/rsi-growth-curves-which-friction-binds]] — both this report and Anthropic's locate the binding constraint *outside* cognition (the slowest un-acceleratable step coupling the loop to reality); the embodied bottleneck re-paces rather than halts, data-wall/research-harder demote into compute, and the abstraction barrier is the one candidate *fundamental* blocker.
  • Multi-Agent Collective Intelligence
    • Do homogeneous LLM collectives produce real synergy, or only humans-with-human-limits benefit from division of labor?
    • What's the actual shape of "multi-agent scaling laws," and does it depend on organization form (homogeneous collective vs. heterogeneous market) or task complexity?
    • Is running more instances more compute-efficient than making individual models larger (up to a single monolithic system)?
    • How do humans meaningfully interact with and steer very large agent groups operating at superhuman speed and output volume?
  • Open-Weight Elicitation Irreversibility
    • **What would an open-weight safety evaluation even report?** A single number is meaningless per premise 1. A curve of dangerous capability against elicitation budget is publishable — and is also a roadmap. Is there a disclosure regime that is informative to auditors and not to attackers?
    • Does the "everybody can audit" advantage actually materialize? Who has funded a serious post-release dangerous-capability audit of any open-weight model, and at what budget?
    • Gemma 4's safety section reports no numbers. Is that a deliberate non-disclosure, a judgment that the model is far from any threshold, or simply a technical report's genre convention? The document does not say, and the distinction matters.
    • Anthropic's answer to a threshold-crossing model was a safeguarded SKU and an unsafeguarded one ([[claude-fable-5]] / Mythos 5), both hosted. What is the open-weight equivalent of shipping the safeguarded SKU?
  • Recursive Self-Improvement
    • The RSI extrapolation rests on trends staying exponential rather than S-curving — but the essay concedes it cannot rule out an architectural ceiling or a compute/energy supply-chain constraint. Which binds first? **Synthesized against DeepMind:** [[wiki/derived/rsi-growth-curves-which-friction-binds]] — the three futures map one-to-one onto DeepMind's three growth shapes; the first friction to bind is the already-binding one (Amdahl's-law verification/oversight = DeepMind's embodied bottleneck), and the abstraction barrier supplies the mechanism Anthropic lacks for whether taste is a *real* ceiling (Future 1).
    • If misalignment compounds through self-improvement (future 3), is AECI-gated [[responsible-scaling-policy-evals|RSP]] review fast enough to catch it before control is lost?
  • Research Taste as the Human Bottleneck
    • How do you measure rubber-stamping? "Humans set direction" can be true on paper while real judgment quietly transfers to the model.
  • Researcher Uplift from Code Output
    • The whole chain rests on **β = 0.5** (pre-AI coding time share), fixed "for simplicity." Kwa flags substantial uncertainty; how much does the 2.3–2.9× band widen once β is varied and measured against Anthropic's actual time-use data?
    • Greenblatt's 0.55/0.45 labor/compute split is itself an assumption. Is the true R&D production function really that insensitive to labor — and if so, does labor uplift matter far less than the RSI discourse assumes?
  • Responsible Scaling Policy Evaluations
    • The two new general-access risk pathways (other AI developers; major governments) are newly in scope but lightly evaluated — what would a positive finding there even look like?
    • How does the RSP brake interact with [[recursive-self-improvement]]: is AECI-based gating fast enough if acceleration compounds, and does single-lab gating even matter without the multilateral [[frontier-pause-verification|pause-verification]] regime?
  • The Abstraction Barrier
    • Is the current paradigm of large-scale pretraining on human data *fundamentally* bounded by human conceptual frameworks, and by how much? (Report open question 1i.)
    • Does the embodied bottleneck reduce the intelligence-growth rate to empirical-science speed, and can that be modelled?
    • Can a system be built that does grounded concept discovery from raw sensor data — and is collective ASI a way around an individual cap?
  • Transformative Creativity
    • Does increasing intelligence inherently produce increasing creativity, or do transformative leaps require something (grounded discovery) the current paradigm lacks?
    • Is the AlphaGo→AlphaFold class strictly exploratory, or are there early signs of transformative (new-conceptual-space) creativity?
    • Could transformative *artistic* creativity ever emerge from optimization power without lived cultural grounding?
  • Universal AI (AIXI)
    • Does modern agentic scaffolding (or RL-tuned implicit decision-making) actually satisfy the AIXI planning ideal, or only superficially resemble it?
    • Can the embedded/multi-agent AIXI extension produce *practical* insight for real multi-agent ASI ([[multi-agent-collective-intelligence]]), or does it remain a theoretical patch?