A tier-list screenshot with no methodology attached is not evidence — it's a conversation starter, and this one started a big one. On September 21, 2026, an X user posting as Sayo shared an image ranking agent harnesses from S to F, captioned simply "ai agent/harness tiers." It reached nearly 80,000 views within hours, and the replies disputed almost every placement in it — starting with the harness that isn't on the list at all.
What the list actually says
| Tier | Harnesses |
|---|---|
| S | Oh My Pi |
| A | Factory Droid, Pi, Zed, fx |
| B | Codex, OpenCode, Cursor, Amp, Devin, Claude Code, Grok Build |
| C | Warp |
| D | Roo Code, Cline |
| F | Antigravity, GitHub Copilot |
No scoring rubric, task set, or usage disclosure accompanied the image — just the ranking itself, dropped without further comment beyond a follow-up post noting a dislike of "the over-use of emoji in default setup" on one entrant.
TL;DR — what people are actually asking
| Question | Direct answer |
|---|---|
| What topped the list? | Oh My Pi, alone in S-tier |
| What's the biggest omission? | Hermes — not ranked anywhere, called out repeatedly in replies |
| What's the most disputed placement? | Claude Code and Codex in B-tier, alongside Cursor and Grok Build |
| Is there a scoring methodology? | No — one commenter asked directly and got no answer |
| Is OMP actually S-tier material? | Disputed — one reply called it "slop" vs. plain pi |
| Should you pick a harness from this list? | No — treat it as one workflow's opinion, build your own eval |
The Hermes omission is the real story
The single most repeated reaction across the replies wasn't about any specific tier placement — it was that Hermes, Nous Research's open agentic harness, doesn't appear on the list at all. At least four separate commenters asked variations of "where Hermes?" independently, without prompting each other, which is a stronger signal of a genuine gap than a single complaint would be. One reply went further, saying they'd switched their entire workflow to Hermes for coding and calling both the Hermes omission and Claude Code/Codex's B-tier placement "beyond dumb" in the same breath.
That reaction matters because Hermes has a real, documented technical profile — a remote-VPS-and-Telegram-CLI-oriented open harness with its own design philosophy distinct from every entrant that did make the list. Omitting an actively-used harness with its own committed user base isn't a minor oversight in a list this specific; it's evidence the ranking reflects one person's own tool exposure rather than a survey of the actual landscape.
Is B-tier for Claude Code and Codex defensible?
This is where the pushback gets more substantive than "you forgot my favorite tool." One reply called grouping Codex and Claude Code into B-tier, below OpenCode in the same tier and below a themed config wrapper in S-tier, "beyond dumb" — not because B-tier is a bad grade in isolation, but because it flattens a real distinction. explainx.ai's own comparison of the leading harnesses treats Claude Code and Codex as full-featured, terminal-first, deeply extensible systems — hooks, skills, subagent orchestration, MCP tool integration — that sit in a genuinely different design category from Cursor, an IDE-integrated product built around a different interaction model entirely. Putting all three in one bucket answers "did I use this today" more than it answers "which harness has the deepest extensibility surface" or "which harness handles a 200-turn autonomous session most reliably" — different questions that would plausibly produce different rankings.
Devin, also placed in B-tier, is a fully autonomous agent product rather than a developer-driven harness in the same sense as the terminal tools around it — another example of the list treating meaningfully different product categories as directly comparable on one axis.
Is Oh My Pi actually S-tier material?
The single sharpest technical disagreement in the thread was about the top slot itself. Oh My Pi (OMP) — a themed configuration layer built on top of Mario Zechner's minimal pi harness, following the same naming convention as Oh My Zsh for shell configs — got one blunt reply: "OMP is slop unfortunately stock pi with bare minimum for your workflow is much better." Sayo's own response defended it while separately admitting a dislike for OMP's default emoji usage, which is a strange combination of endorsing the S-tier ranking while criticizing one of its defaults in the same reply.
The underlying disagreement is a real one in harness design more broadly: does a themed, opinionated configuration layer on top of a minimal harness represent genuine added value, or does it just add surface area over a tool that was already good specifically because it shipped with a small, unopinionated core? Pi's own design philosophy — no baked-in MCP, no default sub-agents, extend only what you need — is arguably in tension with a themed wrapper being ranked a full tier above the base tool it's built on.
The tiers nobody argued about — and why that's suspicious too
It's worth noting what didn't draw pushback, because silence in a reply thread this contentious is its own signal. Warp sat alone in C-tier without comment. Roo Code and Cline — both mature, widely-used open-source VS Code extension agents — landed in D-tier with no defense mounted for them, despite both having real adoption in the wild. Antigravity and GitHub Copilot in F-tier, the list's harshest grade, also went unchallenged in the visible replies, even though Copilot in particular has one of the largest install bases of any tool on the entire list by a wide margin.
That's a pattern worth sitting with: the placements that generated real argument were the ones near the top (S and A tier) and the one most people recognized as clearly wrong (Hermes's absence), while the bottom two tiers passed without a single visible defender. Either the community broadly agrees Antigravity and Copilot deserve F-tier — a real possibility, since Copilot's agentic capabilities have historically trailed purpose-built coding-agent harnesses even as its install base dwarfs them — or the people who'd disagree simply didn't see the post or didn't bother replying. A tier list's uncontested placements aren't automatically correct; they're just placements nobody happened to push back on in this specific thread, which is a different and weaker form of validation than it looks like at a glance.
No methodology, no rank — and someone asked directly
Perhaps the cleanest single critique in the replies came from a commenter who simply asked Sayo: "this is based on what? What is your rank?" That question — never answered in the visible thread — is the actual problem with any viral tier list like this one. A ranking with no disclosed task set, no usage volume per harness, and no scoring rubric is indistinguishable from a personal preference dressed up as a verdict. Another reply put it more bluntly: "No way you have used all of these harnesses intimately enough to put them in tiers."
That's not a reason to dismiss the list — Sayo's bio describes them as a "software factory manager," a role that plausibly involves real exposure to several of these tools — but it is a reason to treat it as one practitioner's dashboard, not a benchmark result. It doesn't disclose what tasks were run, on what codebases, at what volume, or over what time period, which are exactly the variables that would make one harness outperform another for a specific reader's actual workflow.
What to do instead of copying someone else's tier list
- Build your own ranking from your own repository and tasks. A harness that excels at long autonomous refactors may lose badly on quick interactive fixes, and vice versa — the same discipline covered in how to read AI benchmarks applies directly to harness comparisons, not just model benchmarks.
- Separate "harness design philosophy" from "which one I happened to open today." Terminal-first extensible harnesses (Claude Code, Codex, OpenCode), IDE-integrated products (Cursor, Zed), fully autonomous agents (Devin), and minimal ownable cores (pi) are different categories solving different problems — a single ordinal ranking across all of them answers a narrower question than it appears to.
- Check whether a harness you rely on is simply missing before trusting a list's floor as much as its ceiling. Hermes's total absence here is the clearest example of why a tier list's silence about a tool tells you more about the list's author than about the tool.
- Use a structured comparison when you actually need to decide. explainx.ai's top 10 harnesses guide breaks the same landscape down by license, extensibility, and workflow fit rather than collapsing it into one axis.
Related on explainx.ai
- Top 10 Closed-Source and Open-Source Agent Harnesses (2026)
- Pi Agent Harness: Mario Zechner's Minimal Coding Agent You Can Own
- Hermes Agent: Nous Research's Remote VPS + Telegram CLI Guide
- What Is Harness Engineering for AI Agents?
- How to Read AI Benchmarks
- What Is an Agent Harness? Complete Guide
- AgentRun: How Grep.ai Turns Agents Into Cheap, Auditable Workflows
This post reflects a viral X tier-list image and its replies as of September 21-22, 2026. Tier placements, reply quotes, and view counts are as they appeared on X at time of writing; the list's author disclosed no scoring methodology, and this post does not attempt to independently verify or replicate the ranking. Follow @explainx_ai for updates.
