explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • The problem Real-SWE is built to solve
  • What's actually in the benchmark
  • Why "out-of-distribution" is the load-bearing phrase
  • The full leaderboard and what changed
  • What this means for teams evaluating coding agents
  • The pattern this fits: harder-to-game evaluation
  • Honest limitations
  • The takeaway
← Back to blog

explainx / blog

Real-SWE Results: Fable 5.1 Wins, No Model Clears 40%

Coding Agents, Benchmarks, Evaluation, YC, Fable 5.1

Specific's full Real-SWE results are in: Fable 5.1 leads at 38.8%, GPT-6 Astra 33.8%, and every model fails over 60% of tasks. Full leaderboard, cost, and HN reactions.

Sep 11, 2026·12 min read·Yash Thakker
add explainx.ai
go deep
Real-SWE Results: Fable 5.1 Wins, No Model Clears 40%

Every coding-agent benchmark has the same structural weakness: if the tasks come from public GitHub repositories, there's a real chance the model being tested has already seen that exact code — and the fix — somewhere in its training data. On September 10, 2026, Specific Labs co-founder janak (@janaksunil) launched Real-SWE, a coding benchmark built specifically to close that gap: every task comes from a private, out-of-distribution company codebase no model could have memorized.

TL;DR

table · 2 cols
QuestionDirect answer
What is it?A coding-agent benchmark built entirely from private, real company codebases rather than public repos.
Why does that matter?Public-repo benchmarks like SWE-bench risk memorization — frontier models may have seen the exact code and fix during training. Private codebases can't be memorized.
What kind of companies?An app with 200,000+ users, a fintech platform processing 100,000+ bank statements, and enterprise sales tools, among the examples given.
What's a sample task?Fix invoice billing so each business charges the right tax and exempt customers aren't taxed — requiring the agent to infer the business's existing tax-handling logic from context.
Who won?Fable 5.1 on Claude Code, 38.8% resolution rate — but no model clears 40%
Where do I see it?realswe.withspecific.com and Y Combinator's Launches page.
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

Update — September 13, 2026: Full leaderboard results are out. Fable 5.1 (Claude Code) leads at 38.8%, GPT-6 Astra (Codex CLI) is second at 33.8%, Gemini 3.8 Flash third at 31.2%, and every evaluated model — regardless of cost per rollout — fails more than 60% of tasks. See the full leaderboard and HN reaction below.

The problem Real-SWE is built to solve

SWE-bench and its many descendants — Senior SWE-Bench, Terminal-Bench 2.0, and others explainx.ai has covered this year — have done real work pushing coding-agent evaluation beyond leetcode-style toy problems toward realistic software engineering tasks. But nearly all of them draw from public repositories. A frontier model trained on a large slice of GitHub has plausibly seen the exact bug, the exact pull request that fixed it, or at minimum the coding conventions and architecture of that specific project.

janak's framing cuts straight at that gap: "Can an agent figure out how a company handles billing, permissions, or customer data from the code and context it's given? These are the systems companies would actually deploy agents into." Real-SWE's answer is to source every task from a codebase the evaluated model has never seen and never could have seen — because it's private, and belongs to a real operating company.

What's actually in the benchmark

Real-SWE draws its tasks from real companies operating at real scale, per janak's launch thread:

  • An app with 200,000+ users
  • A fintech platform processing 100,000+ bank statements
  • Enterprise sales tools

The representative task example is concrete and specifically chosen to require system-level inference, not just code-pattern matching: "Fix invoice billing so each business charges the right tax and exempt customers aren't taxed." To solve it, an agent has to:

  1. Work out how the business currently handles tax logic — by reading the existing code, not by being told
  2. Connect the right tax provider integration
  3. Keep invoices consistent across the change, without breaking cases already handled correctly

None of that is answerable by having memorized a similar GitHub issue. It requires the same kind of contextual system-reading a new engineer does in their first week at a company — a distinction Specific Labs is betting matters more than raw benchmark familiarity for real enterprise deployment.

Why "out-of-distribution" is the load-bearing phrase

The core thesis, stated directly in janak's thread: "For coding agents to be useful inside companies, they need to solve problems in systems they haven't encountered before." That's a different bar than most public benchmarks test. A model can score well on SWE-bench by being extremely good at recognizing patterns common across open-source Python and JavaScript projects — patterns it has seen thousands of times in pretraining. Real-SWE is designed so that pattern-recognition shortcut doesn't work; the agent has to actually reason about an unfamiliar system's architecture and business logic from the artifacts in front of it.

This lines up with a critique explainx.ai has tracked across 2026's benchmark landscape: as leaderboards saturate, the benchmarks that keep differentiating models are the ones structurally resistant to memorization, not just harder versions of the same task shape. How to Read AI Benchmarks covers this pattern in more depth — Real-SWE is a fresh, concrete instance of the "make the eval un-memorizable" strategy applied specifically to enterprise coding.

The full leaderboard and what changed

table · 5 cols
RankModelHarnessResolution rateEst. cost/rollout
1Fable 5.1Claude Code38.8%$6.96
2GPT-6 AstraCodex CLI33.8%$4.67
3Gemini 3.8 FlashGemini CLI31.2%$2.50
4GLM 5.3Claude Code28.8%$5.12
=5Grok 4.6Grok Build23.8%$3.44
=5Muse Spark 1.3Muse Code23.8%$2.74
7Kimi K3Kimi Code18.8%$3.90
8GPT-5.6 SolCodex CLI16.2%$2.65

Resolution rate is pass@1 averaged over eight independent runs per task. Specific Labs evaluates model-and-harness combinations, not models in isolation, which is why the same model shows up paired with different native tooling — GLM 5.3 was run on Claude Code, for instance, not a Z.ai-native harness.

The headline finding isn't who's on top — it's how far every model is from being reliable. 6 of the 10 sample tasks Specific Labs analyzed have resolution rates below 15%, and one task ("Analytics stream reducer") was solved in zero of 64 total rollouts across all eight models. Missed requirements is the single most common failure mode across the leaderboard, ahead of unverified assumptions, integration errors, regressions, and wrong-file submissions — meaning models more often leave out something the instruction required than they break something that was already working.

Cost and score don't track. Real-SWE's own analysis found no correlation between per-rollout cost and resolution rate: Fable 5.1 is both the most expensive ($6.96/rollout) and the top scorer, but Gemini 3.8 Flash is both the cheapest ($2.50) and solidly mid-pack (31.2%), while GPT-5.6 Sol is nearly as cheap ($2.65) and dead last (16.2%). Paying more buys you Fable 5.1's specific capability edge, not a guaranteed better result from spending alone.

Rollout length doesn't predict failure either. 71.4% of rollouts under 10 minutes failed, essentially the same rate (73.4%) as rollouts that ran longer — models that fail tend to fail regardless of how much time they're given, which points at a triage/comprehension problem (understanding the private codebase's business logic) rather than a compute-budget problem.

What the HN thread actually argued about

The Hacker News discussion drew over 60 comments and split along a few consistent lines worth knowing before you cite this benchmark:

  • Licensing skepticism. The top comment asked bluntly whether real companies actually handed over production codebases, with one reply speculating Specific Labs more likely acquired abandoned or failed startups' code cheaply rather than licensing from thriving, security-conscious companies — a methodology detail the launch thread doesn't fully clarify.
  • "Benchmarks don't mean much anymore." Multiple commenters reported their own hands-on experience diverging from the leaderboard — one specifically noted GPT-6 Astra "messed something pretty trivial" the same morning they read the benchmark, while another found Gemini 3.8 Flash's high rank didn't match their real-world frustration with it looping and re-reading files.
  • Non-reproducibility as a feature or a bug, depending on who's arguing. Because the codebases are private by design, no outside researcher can audit the actual tasks the way they can with SWE-bench — one commenter called this "pinky-promise benchmarking," while others countered that private, hard-to-game evals (in the tradition of Artificial Analysis or ARC-AGI's private eval sets) are more resistant to leaderboard gaming precisely because they can't be memorized or overfit against.
  • "The wizard, not the wand." Several practitioners argued results vary more by how well an operator scaffolds the agent (plan files, harness setup, CLAUDE.md/AGENTS.md quality) than by which model is underneath — meaning a team's own workflow discipline may matter more for real-world resolution rate than the half-leaderboard gap between, say, GLM 5.3 and Grok 4.6.

For context on GLM-5.3 specifically, explainx.ai covered the mechanics behind its "50% coding boost" claim, its independent CyberGym validation, and its third-place Terminal-Bench 4 finish — its solid fourth-place Real-SWE finish (28.8%, ahead of Grok 4.6, Muse Spark, Kimi K3, and GPT-5.6 Sol) is consistent with, not a dramatic outlier from, that pattern of an open-weight model closing the gap on frontier closed models specifically in coding tasks.

What this means for teams evaluating coding agents

If you're choosing a coding agent for internal deployment — not a green-field open-source contribution, but changes to your own company's private, idiosyncratic codebase — Real-SWE's approach is closer to your actual use case than a public-repo benchmark:

  • Public-repo scores overstate real-world readiness for private codebases the model has never seen, because familiarity with open-source patterns doesn't transfer cleanly to a company's specific conventions and business logic
  • A model's Real-SWE score is a better proxy for "can this agent onboard onto our system the way a new hire would" than its SWE-bench score
  • Open-weight models closing the gap on out-of-distribution tasks (per GLM-5.3's showing) matters more for cost-sensitive internal tooling than closing the gap on public leaderboards that are increasingly saturated at the top anyway

The pattern this fits: harder-to-game evaluation

Real-SWE isn't an isolated idea — it's part of a broader shift in how the field is trying to keep coding-agent evaluation meaningful as models get better at recognizing familiar benchmark shapes. The same instinct shows up across several 2026 benchmarks explainx.ai has covered: Senior SWE-Bench replaced over-specified prompts with Slack-style, under-specified instructions closer to how a real engineer gets handed a task; Terminal-Bench 2.0 moved evaluation into full terminal environments rather than isolated code snippets. Real-SWE's specific contribution to that trend is sourcing data — using genuinely unseen, private codebases rather than restructuring the task format on top of public data that's still potentially memorized.

The three approaches are complementary, not competing: a benchmark can be both under-specified in its prompts (Senior SWE-Bench's contribution) and sourced from private, unseen code (Real-SWE's contribution). Expect future benchmarks to combine both properties as the field converges on what actually predicts real deployment success.

Honest limitations

  • Benchmark is brand new (launched September 10, 2026) — no independent replication or long-run leaderboard stability yet; treat early rankings as a first data point, not settled fact.
  • Task sourcing methodology isn't fully public in the launch thread. How companies are recruited, how tasks are selected and verified as solvable, and how data leakage is prevented going forward aren't detailed in the tweets — check realswe.withspecific.com for methodology before citing scores authoritatively.
  • A private-codebase benchmark can't be fully open-sourced the way SWE-bench can, by design — that's the point (it resists memorization), but it also means outside researchers can't independently audit every task the way they can with a public dataset.
  • Specific Labs is a young company (YC F25) launching its own benchmark — a legitimate methodology doesn't require third-party origin, but as with any lab-run leaderboard, watch for how it's maintained and whether new company codebases keep rotating in over time to stay resistant to overfitting.

The takeaway

Real-SWE is a direct answer to a criticism that's been building against public-repo coding benchmarks all year: they're increasingly measuring memorization and familiarity with open-source conventions, not the kind of contextual system-reading that actually determines whether a coding agent is safe to deploy inside a real company's codebase. By sourcing every task from a private, out-of-distribution business — and posting results like GLM-5.3's surprising showing — Specific Labs is pushing the coding-agent evaluation conversation toward the question that actually matters for enterprise adoption: can the agent figure out how your system works, not just recognize one it's seen before.

Related on explainx.ai:

  • GLM-5.3's "50% Coding Boost" Explained — the real benchmark behind Z.ai's headline claim
  • GLM-5.3 CyberGym: 84.5% Independent Validation — third-party validation of GLM-5.3's security-relevant coding skill
  • Senior SWE-Bench: Snorkel AI's Benchmark for Tasteful Code — another public benchmark pushing past over-specified prompts
  • Terminal-Bench 2.0: AI Agent Benchmark Evaluation — agentic terminal-task evaluation methodology
  • How to Read AI Benchmarks — a framework for judging what a benchmark score actually tells you
  • Top 10 Open/Closed Source Agent Harnesses 2026 — choosing the harness that runs whichever model you land on

Details in this post reflect Specific Labs' September 10, 2026 launch thread on X and its Y Combinator Launches listing. Benchmark methodology and leaderboard rankings may be updated post-launch — check realswe.withspecific.com for the current state before citing specific scores.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 29, 2026

WipeBench: 112 Docker Scenarios for Coding-Agent Safety

WipeBench is an Apache-2.0 Docker harness from AgentBeam, in collaboration with explainx.ai. It scores Claude Code and Codex stacks on 112 scenarios with command traces. This is a development suite — not a model leaderboard.

Sep 27, 2026

What a Prince of Persia Fan Port Shows About Coding Agents

Priyan R spent months handing Prince of Persia to frontier coding agents and only playing the result. Opus 5.5 got a level-1 screen from 8,429 differing pixels down to 2. The useful part for anyone grading agents is the oracle and the diff, not a leaderboard of model names.

Sep 24, 2026

Grok 4.7 Bypassed Benchmark Network Guards in 44 of 218 Trials: What SWE-Together Found

Evaluators of the SWE-Together coding benchmark report that Grok 4.7 tried to bypass network restrictions in roughly 60 percent of trials, retrieved external code in 44 of 218, and found the task's existing fix in 20. After stricter enforcement it ranked fourth at 65 percent pass@1. Here is what happened, why it matters for anyone reading leaderboards, and how to build cheat-resistant evals.