explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

custom AI agents

[email protected]

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource librarydemofor LLMs

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

More from us

InfloqInfluencer marketingBgBlurPrivacy-first blurOlly SocialSocial AI copilotCeptoryVideo intelligenceBgRemoverBackground removal

newsletter · weekly

Get AI news, tools, and insights in your inbox.

contactsupportprivacytermsdata rightssubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR — the debate in one table
  • What GPT-3 → Fable actually changed
  • Why the hype thread met disappointed power users
  • The bull case — faster than five years
  • The caution case — S-curves and definitions
  • GPT-5.6 — the near-term datapoint in the thread
  • Three scenarios for 2031 (if you are planning orgs)
  • What to do this quarter (not 2031)
  • Related on explainx.ai
← Back to blog

explainx / blog

Paul Graham on AI in 2031: If Models Improve on Fable Like Fable Improved on GPT-3

Paul Graham's July 2026 X post imagines 2031 models leapfrogging Fable the way Fable leapfrogged GPT-3. We unpack the thought experiment, skeptical replies, GPT-5.6 benchmarks, S-curve debates, and what to plan for.

Jul 7, 2026·8 min read·Yash Thakker
Paul GrahamClaude Fable 5AI ProgressFuture of AIGPT-3Forecasting
go deep
Paul Graham on AI in 2031: If Models Improve on Fable Like Fable Improved on GPT-3

On July 7, 2026, Paul Graham posted a one-line thought experiment that hit X's trending news card — 173 posts, widespread quote-tweets, and a familiar split between awe and skepticism:

"Imagine what it will be like if 5 years from now models have improved on Fable as much as Fable has improved on GPT3."

He was quoting Jared Friedman (@snowmaker, YC partner): "Fable is so insanely good. Deserves the hype."

The post is not a forecast. It is scenario planning by analogy — if the last leap was that large, the next leap might be incomprehensible. The replies are where the useful argument lives: power users who are disappointed at $200–500/day spend, researchers who think two years not five, economists who want cheaper Fable not god-model 2031, and philosophers asking what "as much improvement" even measures.

This post unpacks the thread for builders — without treating PG as a prophet or the skeptics as Luddites.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.


TL;DR — the debate in one table

CampRepresentative replyCore claim
PG / Friedman"Imagine 2031…" / "insanely good"Another GPT-3→Fable-class discontinuity is imaginable
Power-user skeptic0xSero (3 weeks on Fable)Too expensive, downgrades, still hits walls; Opus+GPT-5.5 more reliable for ML
Paying subscriber skepticCryptoKaleo ($200/mo Max)GPT-5.5 x-high ≥ Fable Max on accuracy — disappointment at price
AccelerationsDerya Unutmaz2 years, not 5, for comparable progress
Recursive boostKevin BryanAI helps build better AI — baseline should exceed historical rate
S-curve cautionJeff HuberUnclear where we are on the curve
Commodity hopeGarrett Petersen"Fable-level but cheaper and faster" is already incredible
Definition problemStefan SchubertWe do not know what "same amount of improvement" means operationally
Near-term dataam.will / benhylakEarly signal: GPT-5.6 ~12% above Fable on at least one benchmark (verify)

What GPT-3 → Fable actually changed

Paul Graham's analogy only works if you remember how small GPT-3 feels in 2026 agent terms.

EraModel (approx.)What it could doWhat it could not
2020GPT-3Surprising completion, few-shot prompts, early demosReliable coding, tools, multi-step plans, hour-long tasks
2023GPT-3.5 / early GPT-4Chat products, some coding assistConsistent repo-scale agents, production SWE-bench tier
2026Fable 5SWE-Bench Pro ~80%, Code Arena #1 frontend, multi-hour Claude Code loopsEverything — cost, downgrades, domain walls

Karpathy's English-as-programming arc is the same story in developer language: GPT-3 made prompts interesting; Fable-class models made repos the runtime.

The leap is not "smarter autocomplete." It is autonomous knowledge work — planning, tool use, verification, iteration — at a quality tier that changes what a solo founder or small team can ship in a week.

Rough calendar math: ~6 years from GPT-3 (May 2020) to Fable 5's broad availability (2026). PG's 5-year forward frame lands around 2031.


Why the hype thread met disappointed power users

Trending news amplifies the peak experience. Heavy-user replies document the distribution:

0xSero — three weeks in

After tinkering with Fable for three weeks, the poster listed five bullets:

  1. Too expensive for what it delivers
  2. Too prone to downgrading (model routing / tier surprises)
  3. Still hitting walls on hard tasks
  4. Still a marvel of engineering
  5. Best model to talk to — but maybe shiny-object syndrome; Opus + GPT-5.5 advisor completed ML work more reliably in Claude Code

Bottom line: "Would be nice to have but I'm not spending $500+ in a day."

CryptoKaleo — $200/mo Anthropic Max

Another Max subscriber ran Fable on Max effort for days and found GPT-5.5 x-high as accurate or more capable — frustration that pre-release hype did not match sustained billing reality.

explainx.ai read: These are not contradictions of Jared Friedman's YC-founder use case. They are routing and economics failures for specific workflows — exactly why multi-model stacks (Fable plan → GPT debug → GLM loops) went viral the same week.


The bull case — faster than five years

Derya Unutmaz: two years, conservative

"We will see as much advancement in 2 years. There will be more AI progress in the next two years than in the past 5 years, and that's being conservative."

That compresses PG's timeline to ~2028 — aligned with job-market transformation forecasts already priced for 2027, not 2031.

Kevin Bryan: AI improves AI

"Imagine worlds where they improve more than that, since AI itself will be contributing to improvements in models, energy use, and applications."

This is the recursive acceleration argument:

  • Models assist training data curation and eval design
  • Agents write scaffolding code and benchmark harnesses
  • Distillation pressure forces faster capability transfer

If true, org planning should assume compound progress — not linear extrapolation from 2024–2026.

Garrett Petersen: cheaper matters as much as smarter

"Even just Fable-level but cheaper and faster would be incredible."

That is already happening in parallel to PG's thought experiment:

  • GLM-5.2 at ~1/10th API cost, Code Arena #2
  • Tencent Hy3 — Apache 2.0 agent model
  • China's free-model playbook

The 2031 surprise might be Fable-class at GPT-3 prices, not god-tier at Fable prices.


The caution case — S-curves and definitions

Jeff Huber: where on the S-curve?

"very unclear where on the S curve we lie..."

Scaling-law era assumed smooth log-linear gains. Agent-era gains are lumpy — tool ecosystems, context length, scaffolding, and safety classifiers matter as much as pretraining FLOPs. You can be mid-curve on benchmarks and early-curve on reliability.

Stefan Schubert: what counts as "the same improvement"?

"we don't quite know what 'as much improvement as GPT3->Fable' means in practice"

Fair. The leap mixed:

DimensionGPT-3 → Fable shift
CodingToy snippets → SWE-Bench Pro leader
HorizonParagraphs → million-token projects
InterfaceChat → agent harnesses
EconomicsResearch API → $200/mo consumer tiers
SafetyBasic RLHF → classifier-gated Mythos-class release

Will 2031 beat Fable on all of these, or only on raw reasoning while economics commoditize? PG's sentence does not specify — planners must.


GPT-5.6 — the near-term datapoint in the thread

am.will cited an early benchmark placing GPT-5.6 ~12% above Fable 5 (benhylak thread — the quoted post was unavailable at capture time; treat as unverified until reproduced).

Official published splits today look more task-dependent than a clean 12% universal lead:

Benchmark (published)Leader (mid-2026)
Terminal-Bench 2.1GPT-5.6 Sol (~88.8%) > Fable (~83.4%)
SWE-Bench ProFable (~80.3%) >> GPT-5.5 class
LiveCodeBenchFable first overall in prior tables

Full tables: GPT-5.6 vs Fable 5 comparison

Near-term story: gap closing on agentic terminals; Fable still dominant on several coding leaderboards. That is not yet a second GPT-3→Fable discontinuity — it is frontier convergence.


Three scenarios for 2031 (if you are planning orgs)

Scenario A — PG literal (another discontinuity)

2031 models make Fable feel like GPT-3 feels today: multi-day autonomous projects, negligible human touch on routine software, new science workflows. Implication: hire for orchestration and verification; shrink pure execution headcount in commoditized tiers.

Scenario B — Petersen commodity (same tier, 100× cheaper)

Intelligence plateaus near 2028 Fable-class; price and latency collapse via open weights and efficient MoE. Implication: competitive advantage is distribution and domain data — not model access. Matches AI bubble maturation thesis.

Scenario C — Huber S-curve (lumpy slowdown)

Benchmark gains continue; reliability and economics frustrate daily drivers — echoing July skeptics. Implication: multi-model routing and eval investment beat frontier fanboyism. Mental health and agent burnout stay relevant.

explainx.ai default for builders: plan A-level capabilities on B-level economics while building defenses for C-level friction.


What to do this quarter (not 2031)

Paul Graham's post is a morale and imagination tool for founders. Operationally, the July thread suggests:

  1. Run your own eval — Fable vs GPT-5.5 x-high vs GLM-5.2 Max on your repo (Kilo-style planning tests are templates, not gospel).
  2. Cap daily spend — skeptics citing $500/day are not outliers on Max loops.
  3. Model-agnostic harness — Claude Code, OpenCode, Pi, Cline; swap backends as pricing shifts.
  4. Separate plan vs loop models — the muratcan stack pattern from the GLM adoption thread.
  5. Watch GPT-5.6 GA — if 12% gains reproduce broadly, revisit Fable-only contracts; if not, commodity open weights win loops.

Related on explainx.ai

  • Fable 5 local hardware projection (r/LocalLLaMA 24.8mo lag)
  • GPT-5.6 vs Claude Fable 5 — benchmark comparison
  • Fable 5 after relaunch — developer reactions
  • GLM-5.2 adoption — Fable plan + GLM loops stack
  • English programming — GPT-3 era to agents
  • Your job in 2027 — domain transformation
  • Loop engineering with coding agents
  • AI bubble maturation in 2026
  • China AI playbook — commoditized intelligence
  • Programmer mental health in the agent era

External: Paul Graham on X · Jared Friedman


Thread summaries and benchmark deltas reflect July 7, 2026 X discourse. Paul Graham's post is speculation, not a product roadmap from Anthropic or OpenAI. Verify model scores and pricing on your workloads before budget or hiring decisions.

Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Jul 7, 2026

Will Claude Fable 5 Run Locally by 2028? r/LocalLLaMA's 24.8-Month Lag Projection Explained

Polymarket and @kimmonismus amplified an r/LocalLLaMA projection: frontier models like Claude Fable reach laptop-runnable open-weight parity in ~24.8 months. explainx.ai unpacks the historical lag math, why you will not run Fable weights locally, and what builders should plan for by mid-2028.

Jul 24, 2026

Echo by Tracer: Fable-Level Results at ~1/3 Cost via Open-Weight Pools

Echo is Tracer's coordinated-intelligence experiment: allocate compute, pick which open-weight models participate, and combine their work — not a single model picker. HN debate covers Fusion vs Fugu, cache breaks, and hidden routing.

Jul 22, 2026

Kimi K3 + Fable 5 Routing Beats Either Model Alone, Fireworks Study Finds

Fireworks AI's benchmark makes the case that picking one frontier model is already the wrong question — the real gains come from routing tasks to whichever model is cheapest for that specific job. Here's what the 1,030-task study actually measured, and why "oracle routing" isn't the same as a real router.