Update — August 14, 2026: Pathway shipped a direct successor to this line of research — BDH-CQ, a 150M-parameter "post-Transformer" model claiming 11x cheaper ARC-AGI-1 reasoning than GPT-5.6 Luna using the same latent-space-over-token-space reasoning approach described below.
Most AI scaling conversations still default to one strategy: bigger models, more data, longer context. HRM and TRM added a different axis in 2025: more recursive computation at inference time without proportionally increasing parameter count.
This post summarizes the key ideas from recent HRM/TRM research and the Decoded discussion transcript you shared, then maps those ideas to practical model-design choices.
TL;DR

| Question | Short answer |
|---|---|
| What changed? | HRM/TRM showed that small models can gain strong reasoning behavior via recursive latent refinement loops. |
| Why was it interesting? | Reported ARC-style results were strong relative to model size and training data. |
| Core mechanism | Reuse the same weights repeatedly over internal states (z, z_low, or equivalent) at inference and training. |
| HRM key idea | Two-timescale recursion (high-level + low-level modules). |
| TRM key idea | Simplify to one tiny shared network and keep recursive refinement + deep supervision. |
| Big takeaway | Inference-time recursion is a meaningful compute axis, not just parameter scaling. |
Primary sources:
- HRM paper: arXiv:2506.21734
- TRM paper (“Less is More”): arXiv:2510.04871
- ARC benchmark context: ARC Prize
Why Recursion Re-entered the Conversation
A transformer forward pass is highly parallel and efficient for training, but many reasoning tasks are effectively multi-step algorithms. If the task is hard to compress into one pass, performance can bottleneck even when the model is large.
In the transcript, this is framed as a gap between:
- Token-space iteration (chain-of-thought and tool calls)
- Latent-space iteration (internal recursive state updates)
That distinction matters. Token-space traces are useful, but they are constrained by discrete outputs and supervision artifacts. Latent recursion can keep iterative computation inside a continuous state space.
HRM in One Page
Hierarchical Reasoning Model (HRM) proposes two interacting recurrent modules:
- A high-level module for slower abstract updates
- A low-level module for faster local computation
At a high level, training repeatedly:
- Initializes internal states
- Runs nested recursion loops
- Applies a supervised objective
- Repeats refinement
Reported results in the paper include strong performance on reasoning-heavy tasks (including ARC-style settings) with a relatively small parameter budget and limited training samples.
Reference: Wang et al., 2025
What TRM Kept, What TRM Removed
Tiny Recursive Model (TRM) keeps the core recursive refinement intuition but simplifies architecture and training design.
From the paper’s framing:
- Replace dual-network hierarchy with a single tiny shared network
- Keep recursive latent/output refinement
- Use deep supervision-style training across refinement iterations
The paper reports that this simplified setup outperformed HRM on key ARC-AGI metrics while using fewer parameters.
Reference: Jolicoeur-Martineau, 2025
Chain-of-Thought vs Latent Recursion
A useful way to reason about the difference:
| Approach | Iteration medium | Typical failure mode |
|---|---|---|
| Chain-of-thought | Tokens (external text) | Verbose traces, brittle decomposition, inherited token errors |
| Tool-use loops | Tokens + external API calls | Bounded by tool availability and prior knowledge |
| HRM/TRM recursion | Continuous latent state | Training stability and optimization details become central |
This does not make chain-of-thought obsolete. It reframes it as one recursion interface, not the only one.
Why ARC-Style Tasks Fit This Direction
ARC-style problems emphasize abstraction and stepwise transformation. They are often hard to solve via a single direct mapping from input to output.
Recursive latent refinement is naturally aligned with these tasks because it allows:
- Iterative hypothesis updates
- Intermediate state correction before final output
- More compute depth without proportional parameter growth
That is the core reason these papers attracted attention: not just scoreboards, but a different compute strategy.
Engineering Implications
If you are building reasoning systems, these papers suggest a practical design checklist:
-
Separate model capacity from compute depth Capacity (parameters) and iterative depth (recursion steps) should be tuned independently.
-
Treat recursion loops as first-class hyperparameters Refinement steps, supervision depth, and state-reset behavior can matter as much as width/depth.
-
Benchmark for algorithmic generalization, not only text fluency Include tasks where single-pass pattern matching fails.
-
Expect hybrid architectures General-purpose pretrained models plus compact recursive reasoning heads/modules is a plausible near-term direction.
Limits and Open Questions
Important caveats:
- HRM/TRM are not drop-in replacements for broad conversational LLM products.
- Reported gains are strongest on specific reasoning benchmarks; transfer breadth remains an open question.
- Training dynamics (especially truncated backprop choices and recursion schedules) are still under active study.
- Benchmark-specific optimization risk always exists; cross-domain validation is essential.
Practical Positioning in 2026
The most realistic interpretation is not “small recursive models replace frontier LLMs.”
It is: recursive inference is a complementary scaling law. The field can continue scaling pretrained world models while adding stronger latent recursive computation where algorithmic reasoning is the bottleneck.
That matches where many labs are heading across agent systems and reasoning stacks: combine broad priors with targeted iterative computation.
Update — September 2, 2026: A distinct but related architecture, recurrent depth, made news this week when reports tied it to OpenAI's Astra — see explainx.ai's full explainer on recurrent depth for how it differs from HRM/TRM and why it raises chain-of-thought monitorability concerns.
How to compare recursion with a larger single-pass model
Design the comparison around a compute budget, not only parameter count. A smaller network reused many times can perform more work per answer than its size suggests. Report the refinement schedule, stopping rule, wall-clock latency, and hardware alongside the number of parameters. Otherwise "smaller" can accidentally become a claim about cheaper inference that the experiment did not measure.
For an illustrative grid task, create a baseline that predicts the output directly and a recursive candidate that refines an internal state. Evaluate both on the same held-out examples. Give each model a defined inference budget and record exact-match accuracy separately from runtime. This experiment asks whether additional computation helps under your constraints; it is not a reproduction claim for HRM or TRM.
Then vary recursion depth while keeping weights fixed. If quality improves only up to a point, that point matters operationally. Further loops can consume latency without improving the answer, and a poorly behaved update can even degrade a good intermediate state. A production design needs a reason to stop as well as a reason to iterate.
Use the methodology and evaluation details in the HRM paper and TRM paper to identify what a faithful comparison would need to retain. Their benchmark results should remain attributed to their reported settings.
Test whether the model learned the rule or the presentation
Reasoning datasets can contain shortcuts. Change colors, positions, dimensions, or surface wording while preserving the underlying transformation. A model that succeeds only on the familiar presentation has weaker evidence of abstraction than one that handles these controlled variations.
Keep the training and evaluation split explicit. Repeatedly tuning the schedule against a held-out set can gradually turn that set into development data. Reserve a separate final evaluation and describe the selection process. Small-data experiments are especially sensitive to how examples, augmentations, and task families were chosen.
Inspect failures by category. Some tasks may need longer refinement, others may require missing knowledge, and others may be impossible under the input representation. More recursion addresses only certain bottlenecks. If the answer requires retrieving an external fact, repeatedly refining the same state does not supply that fact.
What this means for an application builder
You do not need to implement a recurrent architecture to learn from these papers. The practical habit is to separate capacity, inference effort, evidence access, and evaluation. A slower answer is useful only when the extra work improves an outcome you care about, and a compact model is useful only when it fits the required workload.
For an agent, measure the whole path: model inference, tool calls, retries, and validation. Latent refinement and external tools solve different problems. The former revisits an internal representation; the latter can obtain information or execute an operation outside the model. Combining them is an architectural choice that still requires application-specific evidence.
Maintain a clear evidence boundary in product claims. A successful puzzle benchmark does not establish dependable code maintenance, safe autonomous action, or broad scientific reasoning. Use the research to choose experiments and ask better questions, then make deployment claims only after measuring the deployment task itself.
Related explainx.ai Reads
- What Is Recurrent Depth? AI Reasoning You Cannot Read
- Pathway's BDH-CQ: 150M model, 11x cheaper reasoning than GPT-5.6
- DeepSeek V4 Flash 0731 scores 89% on ARC-AGI at $0.02/task
- AI Benchmarks in 2026
- LLM Context Window Explained (2026)
- What Are Agent Skills? Complete Guide
- AI Models Hallucinate: Why and How to Catch It
Source Notes
This article is based on:
- Your provided Decoded transcript content
- HRM primary paper: arXiv:2506.21734
- TRM primary paper: arXiv:2510.04871
- ARC benchmark site: arcprize.org
Paper results and benchmark standings can change with revised evaluations, replications, and new benchmark versions. Verify against the latest arXiv revisions and ARC Prize updates.
