explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

custom AI agents

[email protected]

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource librarydemofor LLMs

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

More from us

InfloqInfluencer marketingBgBlurPrivacy-first blurOlly SocialSocial AI copilotCeptoryVideo intelligenceBgRemoverBackground removal

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportprivacytermsdata rightssubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR — what people are asking
  • From Keen to Oak — same religion, different church
  • OaK architecture — what Sutton has been pitching
  • The anti-LLM pitch — imitation vs evaluation
  • 20 watts, trillion parameters — literal or lodestar?
  • Bitter Lesson tension — thread FAQ answered
  • Neolab landscape — where Oak sits in July 2026
  • Who should care
  • Summary
  • Related on explainx.ai
← Back to blog

explainx / blog

Richard Sutton’s Oak Lab — New Algorithms for AGI Beyond Static LLM Training

Turing Award winner Richard Sutton co-founded Oak Lab (Toronto, Jul 2026) with Khurram Javed after Keen Technologies — OaK architecture, continual RL from experience, 20W trillion-parameter goal. explainx.ai maps the bet vs LLMs.

Jul 14, 2026·8 min read·Yash Thakker
Reinforcement LearningAGIRichard SuttonOak LabAI ResearchContinual Learning
go deep
Richard Sutton’s Oak Lab — New Algorithms for AGI Beyond Static LLM Training

The Turing Award winner who wrote the RL textbook just opened a neolab betting against the LLM playbook.

In July 2026, Richard Sutton — 2024 ACM Turing Award laureate, co-author of Reinforcement Learning: An Introduction, and the researcher most associated with temporal-difference learning — co-founded Oak Lab in Toronto with Khurram Javed. The pair left John Carmack's Keen Technologies to pursue what Sutton calls "fundamentally new ideas" for AGI: agents that learn from runtime experience, not curated static datasets.

A July 13 post from @MTSlive (~94K views) amplified the launch; @oaklab_ai published its first research note the same day — Learning from experience instead of curated datasets. explainx.ai maps what OaK is, why Sutton broke from Keen, and how this fits the neolab wave chasing continual learning while frontier labs double down on RLHF-scale alignment.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR — what people are asking

QuestionAnswer
Who?Richard Sutton + Khurram Javed — ex-Keen Technologies (Carmack)
Where?Toronto, Canada — not Alberta (Amii/DeepMind Edmonton hub)
When?July 2026 launch; site updated Jul 13
Thesis?RL from experience · world models · continual learning — not static PT
Architecture?OaK — model-based RL; FC-STOMP abstraction loop
Moonshot?~1T params, real-time learn+plan, ~20W
vs LLMs?Sutton: DL is "weak and inefficient" — needs reworking, not tweaks
Acquihire jokes?Social thread bets Anthropic — no deal announced

From Keen to Oak — same religion, different church

Sutton's departure statement (via X, reported by The Decoder and Office Chai):

Shared with KeenOak-specific
Intelligence from runtime experienceReject current deep learning as sufficient foundation
Reinforcement learning centralNew algorithms — not scaling existing DL stacks
AGI ambitionAlberta Plan / OaK roadmap vs Carmack's engineering sprint

Keen Technologies (Dallas) — Carmack's AGI bet, backers include Tobi Lütke — pursues human-level capability with heavy engineering. Sutton and Javed wanted to rethink learning machinery itself, not only ship faster on today's architectures.

explainx.ai read: This is the second high-profile RL-neolab fork narrative in 2026 — talent leaving well-funded AGI shops when algorithmic philosophy diverges, not only compensation.


OaK architecture — what Sutton has been pitching

Oak Lab is named for the OaK architecture — Sutton's NeurIPS 2025 invited talk and MIT CSAIL Dertouzos lecture (May 2026) outline:

As AI has become a huge industry, to an extent it has lost its way. What is needed to get us back on track to true intelligence? Most of all we need agents that learn continually from their first-person experience.

Three structural bets

  1. Everything learns continually — no frozen pretrained backbone with a thin finetune head
  2. Per-weight step sizes — meta-learned via online cross-validation (Sutton's step-size optimization line)
  3. FC-STOMP abstraction loop — grow structure over time:
    • Feature construction
    • SubTask posed from the feature
    • Option learned to solve it
    • Model of the option (world model chunk)
    • Planning on that model

Prior research on oaklab.ai

Paper / themeYearRelevance
The OaK Architecture: A Vision of SuperIntelligence from Experience2025Manifesto
The Alberta Plan for AI Research2023Multi-year roadmap
Horde — sensorimotor knowledge2011Many parallel value functions
SwiftTD, step-size optimization, columnar-constructive networks2023–24Continual / real-time learning machinery
Learning from experience instead of curated datasetsJul 13, 2026First Oak Lab publication
Event-driven NNs, batch-size-one learningComing soonListed on site

Thread replies praised Sutton's NeurIPS OAK talk — the lab productizes a ** decade-long Alberta Plan**, not a weekend manifesto.


The anti-LLM pitch — imitation vs evaluation

Sutton's June 2026 public framing (via press summaries): generative AI imitates but cannot evaluate its own outputs — blocking discovery.

LLM industry loop (2026)Oak Lab counter
Pretrain on curated web/codeLearn online from agent's own stream
RLHF/DPO on human/AI preferencesReward from environment + internal critique
Frozen weights + periodic retrainContinual weight + step-size updates
World as textWorld models + planning

That does not mean Oak ignores OpenAI's beneficial-trait RL results — it reframes them as patching imitation systems rather than building experience-native agents.

For builders, the actionable split:

  • Ship products today → LLM + RLHF oversight stack
  • Research AGI foundations → ask whether online RL + world models escape dataset ceiling and catastrophic forgetting

20 watts, trillion parameters — literal or lodestar?

Press reports Sutton's long-term goal: an agent with ~1 trillion parameters that learns and plans in real time on ~20 watts.

ContextScale
Human brain~20W — often cited in AGI rhetoric
Single H100~700W — runs models orders smaller than 1T interactively
Frontier LLM trainingMegawatts across clusters

Plain read: The number is a north star for efficiency + continual learning, not a 2026 product spec. Oak Lab is an algorithm lab first — hardware co-design may follow (Keen DNA) but is not the launch headline.


Bitter Lesson tension — thread FAQ answered

Neolab skeptics on X asked: Doesn't Sutton's own Bitter Lesson say scale + general methods beat hand-crafted algorithms?

Bitter Lesson coreOak Lab response (implicit)
General methods winAgree — but static backprop on fixed datasets may not be the winning general method for continual agents
Compute + data scaleOak bets experience streams are the data — efficiency matters (20W)
Human knowledge hurtsOaK builds abstractions online (FC-STOMP) instead of hand-engineering features

Sutton is not anti-scale — he is anti-"scale the wrong paradigm forever." Oak Lab is a neolab bet that 2020s DL is the new 2010s chess engines: impressive, industry-defining, not the final AGI substrate.


Neolab landscape — where Oak sits in July 2026

Lab / threadFocus
Oak LabContinual RL, OaK, experience > datasets
KeenCarmack engineering AGI (Sutton alumni)
Frontier LLM labsScale PT + alignment RL
Physical AILeCun world models / Moravec
Agent harnessesClaude Code, OpenClaw — tools atop LLMs, not new learning laws

Social acquihire jokes (Anthropic buys Oak in months) reflect RL talent scarcity — Sutton is the canonical RL citation — not a disclosed deal.

Toronto vs Alberta: Riley RostovM and others noted Oak is not in Edmonton (Amii / historical Sutton base) or Vancouver — but Canada retains the research gravity Sutton built at University of Alberta for decades.


Who should care

AudienceTakeaway
RL researchersWatch oaklab.ai for batch-size-one / event-driven posts
LLM product teamsOak does not invalidate your stack — it questions 10-year substrate
InvestorsNeolab thesis diversity — not every AGI bet is more transformers
Policy / safetyContinual agents raise new monitoring problems — static eval harnesses may not transfer

Summary

Richard Sutton co-founded Oak Lab in Toronto (July 2026) with Khurram Javed after leaving Keen Technologies — pursuing new AGI algorithms built on reinforcement learning from experience, the OaK architecture, and continual world models, not curated dataset scaling. The lab's first note dropped July 13; Sutton's moonshot remains a ~1T-parameter, ~20W, real-time planning agent. Social hype (~94K views) frames a neolab moment; the substance is a direct challenge to LLM pretrain + RLHF as the final path to intelligence — from the researcher who literally wrote the book on RL.


Related on explainx.ai

Update — July 14, 2026: Same week, Demis Hassabis published a frontier AI governance essay betting AGI arrives in a few years via scaling — the opposite timeline thesis from Oak's algorithm-first path.

  • Scalable oversight — RLHF & Constitutional AI
  • OpenAI beneficial trait RL — good generalizes OOD
  • AI alignment introduction — outer vs inner goals
  • What is fine-tuning — RLHF stage explained
  • Yann LeCun — LLMs, physical agents, continual learning
  • History of AI — DeepMind RL lineage
  • Teaching Claude why — alignment vs capability

Sources: Oak Lab · NeurIPS 2025 — OaK invited talk · MIT CSAIL — OaK lecture May 2026 · The Decoder — Jul 13, 2026 · Office Chai · @MTSlive Jul 13, 2026


Oak Lab claims and Sutton quotes reflect July 2026 public announcements. AGI timelines and hardware targets are aspirational — not verified product commitments.

Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Jul 21, 2026

Sakana AI's 'Diffusing Blame': Training Neural Networks Like Real Neurons

Real neurons are fixed excitatory or inhibitory — standard backprop ignores that and needs a biologically implausible trick to work. Sakana AI's Error Diffusion approach learns without it, scoring 96.7% on MNIST and holding up in reinforcement learning on Ant, Humanoid, and Craftax.

Jun 16, 2026

From AGI to ASI: DeepMind's 57-Page Roadmap for What Comes After Human-Level AI

DeepMind researchers published "From AGI to ASI" on June 10, 2026 — a 57-page investigation into how AI might continue developing after it reaches human level. Four pathways, concrete bottlenecks, and a key insight: the transition may not be a single step change but a series of transformative societal shifts.

Aug 3, 2026

Andrew Ho Leaves OpenAI: RSI Quote, Overvaluation, RL Data Startup

Late July–early August 2026: Andrew Ho exits OpenAI after eight months to sell high-end RL datasets (GeneBench-Pro lineage), tells colleagues to take tender liquidity, and becomes a Polymarket headline over a stated preference for “rapid RSI & human disempowerment.” explainx.ai separates the quote from the business thesis.