explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionaryagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • Why this framework matters right now
  • The four stages, in detail
  • Why nobody has clearly reached stage 4 yet
  • Why the industry keeps collapsing the stages together
  • The practical reading for builders
  • The falsifiability the paper's framing buys you
  • Related reading
← Back to blog

explainx / blog

"The Last AI Built by Humans": What Genuine Recursive Self-Improvement Means

Recursive Self-Improvement, AI Safety, AI Research, AGI, Alignment

A new paper argues real recursive self-improvement means AI choosing its own improvement strategy, not just executing human-designed ones. Here''s the four-stage roadmap and why the framing matters more than the hype.

Sep 14, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
"The Last AI Built by Humans": What Genuine Recursive Self-Improvement Means

A paper posted to alphaXiv this month, "The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement," makes an argument that's easy to state and harder to sit with: most of what gets called "recursive self-improvement" in 2026 doesn't actually clear the bar the term implies. The paper's real contribution isn't a new capability claim — it's a four-stage framework that separates AI getting better at tasks from AI getting better at getting better, and argues only the second is genuine RSI.

TL;DR

table · 2 cols
QuestionAnswer
Core claimGenuine RSI means AI choosing and modifying its own improvement process, not just executing human-designed upgrades faster
The four stages(1) executes human-designed improvements → (2) chooses its own improvement strategy → (3) generates its own learning experiences, adapts from deployment → (4) modifies the mechanisms that create future improvements
Where are current systems?Mostly stage 1–2 — human researchers still design the menu of improvement strategies, even when AI picks from it
What does the title mean?Once a system reaches stage 4, subsequent generations are no longer meaningfully "built by humans" — they're built by the prior generation's self-modified improvement process
How does this compare to industry RSI talk?Looser — Amodei's "Pace the Frontier" essay and most 2026 coverage use RSI to mean any AI-accelerates-AI dynamic, without this paper's stricter, falsifiable stage distinctions
Practical useGives builders and policymakers a way to ask "which stage is this system actually at?" instead of treating all self-improvement claims as equivalent

Why this framework matters right now

2026 has been a year of increasingly loose RSI claims. Dario Amodei's We Must Pace the Frontier essay named recursive self-improvement as his primary concern, saying it's "advancing drastically faster" industry-wide. explainx.ai has separately covered a wave of self-described self-improving systems — Ornith-1.5, NeoHourse-1's recursive self-improvement routing harness, and the self-evolving coding agent harness from a "second writer", along with explainer pieces on what recursive self-improvement actually is. Almost all of that coverage uses RSI as a single, undifferentiated label.

This paper's contribution is to insist that label is doing too much work. A system that gets faster or cheaper at a fixed set of human-chosen optimization targets is meaningfully different — in terms of both capability trajectory and safety risk — from a system that starts choosing what to optimize for in the first place, and different again from one that starts rewriting the process by which it decides what to optimize for. Collapsing all three into "RSI" makes every capability announcement sound like the same kind of event, when they're not.

The four stages, in detail

Stage 1: AI executes human-designed improvements. This is the default state of the field, and it's where most of 2026's "self-improving" coding agents and routing harnesses actually sit. A human team defines the training objective, the evaluation metric, and the update mechanism; the AI system executes that loop, often faster and at greater scale than a human team could manage manually, but the design of the improvement process is entirely human. Faster iteration is not the same as autonomous strategy.

Stage 2: AI chooses its own improvement strategies. Here the system is given a menu of possible improvement approaches — different fine-tuning strategies, different architectural tweaks, different data-curation methods — designed by humans, and the AI selects among them based on its own evaluation of what will work best. This is a real step up in autonomy, but it's still bounded: humans defined the option space even if they didn't pick the specific option.

Stage 3: AI generates its own learning experiences and adapts from deployment. This stage breaks the dependency on human-curated training data entirely. The system creates its own training signal from real-world interaction — closer to what Satya Nadella described in his September 14, 2026 post as a "continuous learning loop/hill climbing machine," covered in explainx.ai's breakdown of Nadella's superintelligence principle — turning deployment feedback into ongoing improvement without a human curating what counts as a good training example.

Stage 4: AI modifies the mechanisms that create future improvements. This is the paper's actual bar for "genuine" RSI, and the reason for the title. At this stage the system isn't just improving its own capabilities — it's changing the process that generates future improvements, meaning each subsequent generation is shaped by a self-modified improvement mechanism rather than a human-engineered one. Once that happens, the paper argues, later generations are no longer, in any meaningful sense, "built by humans" — hence "the last AI built by humans" is whichever system first crosses into stage 4.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

Why nobody has clearly reached stage 4 yet

It's worth being precise about what's not being claimed. The paper isn't asserting that a stage-4 system already exists — it's providing the vocabulary to recognize one if and when it appears, and to notice that current industry claims of "self-improving AI" mostly describe stage 1 and 2 behavior dressed in stage-4 language. Ornith-1.5 and similar open-weight self-improving models explainx.ai has covered generally operate within a human-defined training and evaluation loop — impressive stage 1–2 execution, not stage 4 mechanism modification.

That distinction matters for the safety conversation specifically because Amodei's essay treats RSI as a single accelerating trend worth building institutional guardrails around now — see the embedded evaluator concept he proposes as a check. This paper implicitly argues that the guardrails needed for stage 1–2 systems (auditing training pipelines, checking for benchmark gaming) are qualitatively different from what would be needed if a system actually reached stage 4 (auditing a self-modifying improvement mechanism, which by definition resists straightforward human inspection since the mechanism itself is no longer human-designed).

Why the industry keeps collapsing the stages together

There's a structural reason "self-improving AI" gets used loosely rather than precisely, and it's worth naming: stage 1 and stage 2 systems already produce headline-worthy capability jumps, and a vendor announcing a new routing harness or fine-tuning pipeline has every incentive to describe it in the most dramatic available language. "Our system chooses its own training data" (stage 2–3) and "our system is recursively improving itself" (implying stage 4) sound similar in a press release, but they describe fundamentally different levels of autonomy over the improvement process itself. The paper's stage framework is useful precisely because it gives a reader a checklist to interrogate a claim rather than accept the framing at face value: does the system choose what to optimize, or only how fast to reach a human-set target? Does it generate its own training signal, or consume a human-curated one? Does it change the mechanism deciding future improvements, or only the parameters within a fixed mechanism?

This also explains why safety researchers keep circling back to RSI as the single most important variable to track, even as its definition stays contested. A stage 1 system accelerating within a human-designed loop is bounded by the humans' own understanding of what "improvement" should mean — a meaningful safety backstop, even an imperfect one. A stage 4 system that has started modifying its own improvement mechanism has, by construction, moved past that backstop: the definition of "improvement" it's now optimizing toward may no longer be legible to the humans who built the system that built it. That's the actual reason the paper's title carries the weight it does — not because stage 4 has arrived, but because it names the exact threshold where human oversight of the process, not just the outputs, stops being guaranteed.

The practical reading for builders

If you're building agentic systems, harnesses, or evaluation pipelines rather than debating AGI timelines, the useful takeaway isn't the apocalyptic framing implied by the title — it's the stage framework itself as a diagnostic tool. When someone describes a system as "self-improving," ask which stage they mean:

  • Is it executing a human-defined optimization loop faster (stage 1)? That's most current agentic coding harnesses and fine-tuning pipelines.
  • Is it selecting among human-defined improvement strategies (stage 2)? That's closer to routing/orchestration harnesses that pick which fine-tuning or prompting strategy performs best on a given task.
  • Is it generating its own training signal from live deployment (stage 3)? That's the frontier most enterprise "continuous learning loop" pitches are actually reaching for.
  • Or is it changing the mechanism that decides all of the above (stage 4)? Nothing publicly documented in 2026 clearly meets this bar yet.

Being able to name which stage a specific claim describes is a better tool for evaluating both vendor pitches and safety-relevant capability claims than treating "recursive self-improvement" as one undifferentiated, maximally alarming phrase.

The falsifiability the paper's framing buys you

One underrated virtue of a four-stage ladder over a binary "is it RSI or not" question is that it's actually falsifiable in a way loose industry usage isn't. A claim like "our system is recursively self-improving" is nearly impossible to verify or refute from the outside — it's vague enough to accommodate almost any capability demonstration a lab wants to present as evidence. A claim like "this system independently generated its own training curriculum from deployment feedback without human curation" (stage 3) is a specific, checkable assertion: you can ask what the training data actually was, who selected it, and whether a human review step gated what got included. Similarly, "this system modified its own hyperparameter-selection or architecture-search process without a human specifying the space of allowed modifications" (stage 4) is a concrete technical claim that can be inspected against actual system logs, rather than argued about in the abstract.

That's the paper's most durable contribution, independent of whether its own four-stage taxonomy becomes the field's standard vocabulary going forward. Any framework that turns "is this dangerous self-improving AI" from a rhetorical question into a checklist of specific, independently verifiable technical claims is doing useful work for the safety conversation — regardless of which particular framework wins out. For journalists, policymakers, and builders trying to evaluate the next lab announcement claiming some version of self-improvement, the operative move is simply asking which stage, specifically, is being claimed, and asking for the technical evidence that would distinguish that stage from the one below it.

Related reading

  • What Is Recursive Self-Improvement (RSI) in AI?
  • Dario Amodei Wants to "Pace the Frontier" — Here's the Actual Plan
  • What Is an Intelligence Explosion? Explained
  • Satya Nadella's Superintelligence Principle and Microsoft's MAI Code of Conduct
  • Ornith-1.5: Self-Improving Open-Weight Model
  • NeoHourse-1: Recursive Self-Improvement Routing Harness
  • What Is an Embedded Evaluator in AI Safety?

Official source: alphaxiv.org/abs/2609.11873

This post reflects the alphaXiv paper summary and public discussion as of September 14, 2026. Check the official alphaXiv listing for the full paper and any revisions.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Sep 10, 2026

What Is an Intelligence Explosion? AI Term Explained

"Intelligence explosion" describes a specific mechanism — not just fast AI progress, but a self-reinforcing loop where a system's own intelligence gains let it improve itself again, faster each round. explainx.ai traces the term from I.J. Good's 1965 essay to Paul Christiano's September 2026 warning at OpenAI, and explains what would actually have to be true for one to happen.

Sep 10, 2026

What Is Recursive Self-Improvement (RSI) in AI?

Recursive self-improvement is what happens when an AI system helps build a better version of itself, which then helps build an even better one. explainx.ai breaks down the mechanism, the 4-level ladder researchers use to measure it, and real systems — AIDE², NeoHorse-1 — already climbing it.

Jun 16, 2026

From AGI to ASI: DeepMind's 57-Page Roadmap for What Comes After Human-Level AI

DeepMind researchers published "From AGI to ASI" on June 10, 2026 — a 57-page investigation into how AI might continue developing after it reaches human level. Four pathways, concrete bottlenecks, and a key insight: the transition may not be a single step change but a series of transformative societal shifts.