Update — September 5, 2026: OpenCode shipped a new stealth model, Omen Alpha, paid-only via OpenCode Go at ~$0.20/M input tokens — same anonymous-launch pattern, unconfirmed identity. Full coverage.
Update — August 26, 2026 (evening): Official name is GLM-5.3-Flash — 320B-A18B, MIT, $0.15/$0.50 API. Full launch guide.
Update — August 26, 2026: Z.AI (Zhipu) confirmed to Bloomberg that Ox Alpha is a new GLM-series model and said open weights drop tonight. Ox Alpha hit #1 on OpenRouter, more than doubling DeepSeek usage — community GLM-5.3-Flash theory vindicated. Full timeline: Ox Alpha identity post.
Update — August 25, 2026: By the end of Ox Alpha's first four-day free preview window, community dashboards and trade reporting cite up to ~26 trillion tokens processed on OpenRouter — roughly 2.6× the prior record for a single-model launch (CryptoBriefing reported 11.6T in 72 hours mid-window). Treat headline totals as platform-analytics estimates, not audited revenue; the takeaway for builders is demand shape — Claude Code, Hermes Agent, and Zed drove agentic, high-context burns while the model was $0/$0. See also Ox Alpha top use cases.
OpenRouter added a new stealth model on August 20, 2026: Ox Alpha (stealth/ox-alpha) — positioned as a frontier-class model for efficient coding, sustained agentic work, and production workloads, with a 1M-token context window, multimodal input, tool calling, and $0/$0 pricing during preview.
The release lands one day after AT&T reported 56% coding-cost savings from model routing and in the same week Stripe's investor letter named OpenRouter its largest-ever acquisition — still expected to close in the coming weeks. Ox Alpha is not about that deal; it is a concrete model you can point an agent harness at today, for free, through the same OpenRouter API many teams already use for Fusion panels and cost-aware routing.
TL;DR — what builders are asking
| Question | Direct answer |
|---|---|
| Model ID? | stealth/ox-alpha |
| Released? | August 20, 2026 (per OpenRouter model page) |
| Price? | Free — $0/M input, $0/M output |
| Context / max output? | 1,048,576 tokens in · 131,072 tokens out |
| Modalities? | Text, images, and video in → text out |
| Tool calling? | Yes — supported; OpenRouter reports ~4.45% tool-call error rate (3-day avg) |
| Throughput / latency? | ~50 tok/s and ~2.02s P50 latency (OpenRouter dashboard) |
| Trains on your data? | OpenRouter says prompts/completions are retained but not used for training (this release); OpenCode separately advertises Zero Data Retention for its own Ox Alpha traffic — see below |
| Who made it? | Z.AI → GLM-5.3-Flash — MIT, 320B-A18B (launch) |
| Top agent traffic? | Claude Code (~9.3B tokens), Hermes Agent (~9.0B), plus Oh-My-Pi, DeepSeek Harness, Z Code on the model's apps chart |
| How does it benchmark? | Independent test by developer Ben Davis: 80% on DeepSWE, ahead of Fable (65%) and GPT-5.6 Sol (52%) — not an official OpenRouter or provider number |
| Free anywhere else? | Yes — OpenCode (the open-source coding agent) also routes to Ox Alpha directly, and OpenCode Go made it near-unlimited and free for 6 more days as of Aug 21 evening, not counted against normal Go usage |
What OpenRouter stealth models are
OpenRouter's stealth models are preview releases from third-party providers who stay anonymous during the preview window. OpenRouter routes traffic; it does not develop or own the weights.
The pattern is established:
- Models ship under codenames (
Ox Alpha, past examples include Hunter Alpha, Healer Alpha, Owl Alpha). - They are often free during preview to drive eval traffic.
- Providers historically logged prompts for improvement — Ox Alpha's page makes a narrower claim for this release (more below).
- OpenRouter later reveals some identities — Hunter and Healer turned out to be Xiaomi MiMo models. Ox Alpha was confirmed as Zhipu GLM on August 26, 2026 (Bloomberg) — open weights promised the same night.
Governance falls under OpenRouter's Stealth Model Terms — separate from standard provider terms. Treat stealth previews as time-boxed experiments, not long-term infrastructure commitments.
Ox Alpha specs (verified from OpenRouter)
These numbers come from the Ox Alpha model page as of August 21, 2026:
| Spec | Value |
|---|---|
| Model slug | stealth/ox-alpha |
| Provider label | Stealth (single upstream host) |
| Listed price | $0 / $0 per million tokens |
| Context window | 1,048,576 tokens |
| Max output | 131,072 tokens |
| Input modalities | Text, image, video |
| Output | Text |
| Reasoning | Supported (reasoning parameter documented) |
| Tool calling | Supported |
| Release date | August 20, 2026 |
| Cache hit rate (3d avg) | ~67% |
| Uptime (3d) | 99.99% |
| Availability (3d) | 96.79% |
OpenRouter's description emphasizes long-horizon software engineering, complex reasoning, and workflows that combine text with visual context — language that matches how agent harnesses actually burn tokens: repo maps, screenshots, test logs, and multi-step tool loops.
Who is already using it
OpenRouter publishes app-level token share for each model. On Ox Alpha's first days, the leaderboard is dominated by agent harnesses, not chat UIs:
- Claude Code — Anthropic's agentic coding tool (~9.32B tokens on the model chart)
- Hermes Agent — Nous Research's persistent open-source agent (~8.98B tokens)
- Oh-My-Pi, DeepSeek Harness, and Z Code — also listed among top senders on the model page
That traffic mix is a signal: teams running real coding agents are treating Ox Alpha as a workhorse model, not a playground curiosity. It does not prove quality on your repo — only that production-shaped workloads are flowing.
An independent benchmark: 80% on DeepSWE
Developer Ben Davis ran Ox Alpha through DeepSWE, a coding-agent benchmark, and reported it scoring 80% — ahead of Fable at 65% and GPT-5.6 Sol at 52%. This is not an OpenRouter or provider-published number; treat it the same way you'd treat any single developer's eval run: a useful data point, not a verified leaderboard entry. It does line up with the traffic pattern above — a model agent harnesses are routing real work to tends to also do well on agent-shaped benchmarks like DeepSWE.
Free pricing and the "no training" claim
Two details matter for anyone routing proprietary code through Ox Alpha.
Free — for now
OpenRouter lists Ox Alpha at $0/$0. Stealth previews have historically been free to accumulate feedback, then either graduate to a named model with list pricing or disappear. Budget for the possibility that "free" is a preview subsidy, not a permanent tier — the same caution AI token pricing guides apply: watch usage dashboards and keep fallback routes.
Retention ≠ training (read both lines)
OpenRouter's Ox Alpha banner states:
Prompts and completions for this model are retained by the provider and are not used for training; all other use is governed by the Stealth Model Terms.
That is a meaningful distinction:
| Claim | What it implies | What it does not imply |
|---|---|---|
| Not used for training | Provider says your completions won't fine-tune weights (this release) | Prompts aren't stored |
| Retained by the provider | Logs likely persist on provider infrastructure | You know who the provider is |
| Stealth Model Terms | Other uses (eval, abuse monitoring, legal holds) may still apply | Enterprise ZDR / HIPAA |
For production secrets, an anonymous provider with retained logs is a different risk profile than calling Anthropic or OpenAI under a signed enterprise agreement — even when the sticker price is zero. The Databricks cost post and AT&T routing story both assume internal evals and tiered routing precisely because model substitution is never just a price question.
OpenCode's version: Zero Data Retention, and a "100T tokens/day" capacity claim
Ox Alpha isn't only accessible through OpenRouter. OpenCode — the open-source coding agent whose traffic already shows up among Ox Alpha's top senders above — announced its own direct route to the model on X:
"Ox Alpha (stealth model) is free for the next week — 1M Context, Multi-modal, Zero Data Retention. Generous rate limits, near unlimited usage. We have capacity for 100T tokens per day, let's see what you can do."
Two things are worth flagging rather than repeating at face value:
- "Zero Data Retention" vs. OpenRouter's "retained but not used for training." These are different claims from different parties in the same request chain. OpenRouter's model page describes the upstream provider's retention policy; OpenCode's ZDR claim describes what OpenCode itself does with traffic that flows through its client. Neither statement contradicts the other outright, but they are not interchangeable — a ZDR promise from the client layer doesn't override a retention policy at the model-provider layer underneath it. If retention is a hard requirement for you, verify the full path, not just the layer that says the reassuring thing.
- The 100 trillion tokens/day figure drew immediate skepticism. Theo (t3.gg) reacted publicly: "'We have capacity for 100T tokens per day' — okay who the fuck made this model and where did they get this much compute?" That's a fair question. For scale, 100T tokens/day is in the range of what only the very largest frontier-model providers claim to serve in total, across every model and customer combined — a single free stealth preview claiming that ceiling is either an extraordinary infrastructure story or marketing rounding. Nobody involved has clarified which.
OpenCode Go — OpenCode's hosted/managed offering — extended the same deal on August 21: near-unlimited, completely free access to Ox Alpha for 6 more days, and usage through it does not count against your normal OpenCode Go quota. If you already run OpenCode Go, this is currently the cheapest way to put real repo work through Ox Alpha without touching your regular usage budget.
Who is Ox Alpha? (GLM-5.3-Flash — August 26, 2026)
Official product name: GLM-5.3-Flash. Z.ai's @Zai_org launch post (Aug 26, 2026, 7:42 PM) states it was "previously previewed as Ox Alpha" on OpenCode and OpenRouter.
| Spec | Value |
|---|---|
| Architecture | 320B total / 18B active |
| Context | 1M tokens, natively multimodal |
| License | MIT — Hugging Face |
| API price | $0.15/M in · $0.50/M out · $0.03/M cached |
| Stealth infra | Entire preview on Chinese AI chips, per Z.ai |
| Local inference | SGLang, vLLM, TokenSpeed |
Full benchmarks, architecture, and routing guide: GLM-5.3-Flash launch post. Forensics timeline: Ox Alpha identity post.
How to try Ox Alpha via the OpenRouter API
OpenRouter's API is OpenAI-compatible. Swap the base URL and model slug.
1. API key
export OPENROUTER_API_KEY="sk-or-v1-..."
Create keys in the OpenRouter dashboard.
2. Minimal curl (non-streaming)
curl https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "stealth/ox-alpha",
"messages": [
{"role": "user", "content": "Refactor this function to handle nil inputs safely."}
]
}'
3. Agent harness / Claude Code / Hermes
Point your harness provider at OpenRouter and set the model to stealth/ox-alpha. Optional headers (HTTP-Referer, X-Title) only affect OpenRouter leaderboard attribution — not model behavior.
For tool-heavy loops, enable streaming and watch cumulative latency: OpenRouter reports ~11.6s median end-to-end latency on multi-step agent turns — acceptable for background agents, painful for tight interactive edits unless you route subtasks to faster tiers.
4. Routing pattern (recommended)
Mirror enterprise playbooks from AT&T's LiteLLM gateway:
cheap/open route (Ox Alpha) → internal eval gate → premium model on failure
Ox Alpha is a strong candidate for the cheap tier while it stays free — not a blind default for merges, security reviews, or regulated data.
What people are asking
Is Ox Alpha actually "frontier"?
OpenRouter's marketing calls it frontier-class for efficient coding and agentic work. There is no independent benchmark sheet on the model page — only usage stats and latency. Community threads compare it favorably to some Chinese open models and note fast tool use, but your golden tasks (lint fixes, migrations, test generation) are the only eval that matters.
Will it stay free after Stripe closes the OpenRouter deal?
Unknown. Stripe's acquisition is not closed yet; Ox Alpha pricing is set by the anonymous provider + OpenRouter listing, not by Stripe today. Watch the model page and your invoice — free stealth previews can change overnight.
Ox Alpha vs OpenRouter Fusion
| Ox Alpha | OpenRouter Fusion | |
|---|---|---|
| Shape | Single stealth model | Multi-model panel + judge |
| Cost | Free (preview) | Sum of panel + judge tokens |
| Best for | High-volume agent loops, coding | Research, deliberation, high-stakes synthesis |
| Identity | Anonymous | Named panel models |
They solve different problems. Fusion is for when being wrong is expensive; Ox Alpha is for when token volume is expensive — at least while the preview lasts.
Should I send customer PII or prod secrets?
No — not to an anonymous stealth provider that retains logs. Use Ox Alpha for sanitized repos, open-source work, or spikes where log retention is acceptable. Same rule as any free preview tier.
Bottom line
Ox Alpha is real, free, and already carrying billions of tokens from Claude Code and Hermes Agent — and Zhipu confirmed on August 26, 2026 that it is a GLM model with open weights releasing that night. The actionable move: add stealth/ox-alpha as a routed tier (or migrate to the named Z.AI slug once OpenRouter updates), run your coding evals, and watch for the Hugging Face drop if you self-host.
Related on explainx.ai
- Omen Alpha: OpenCode's new stealth model at $0.20/M tokens — the sequel stealth launch, September 2026
- Top 10 things people are building with Ox Alpha — fluid sims, 3D scenes, a DeepSWE benchmark run, and more
- GLM-5.3-Flash official launch — Ox Alpha unmasked
- Ox Alpha: forensics timeline
- Why free preview pricing wrecks model evaluation — the corrected Gemini/Ox Alpha timing timeline, plus a GLM-5.3-Flash vs Gemini 3.7 Flash decision table
- Stripe acquires OpenRouter — what builders should know
- AT&T cut AI coding costs 56% with model routers
- Databricks: managing AI coding costs at scale
- OpenRouter Fusion API guide
- Hermes Agent #1 on OpenRouter rankings
- AI token pricing, explained
- Choosing open-weight vs closed models
- What is OpenRouter? Enterprise guide · Loop engineering with coding agents
- Inkling's "free on OpenRouter" headline, decoded — the fine print
Primary sources: OpenRouter — Ox Alpha model page · OpenRouter — Stealth provider · OpenRouter — Stealth Model Terms
Accurate as of August 26, 2026. Zhipu confirmed GLM lineage to Bloomberg; open weights promised Aug 26 night. Stealth slug, pricing, and limits may change after reveal — see identity timeline.
