explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionaryagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR — what builders are asking
  • What OpenRouter stealth models are
  • Ox Alpha specs (verified from OpenRouter)
  • Free pricing and the "no training" claim
  • Who is Ox Alpha? (GLM-5.3-Flash — August 26, 2026)
  • How to try Ox Alpha via the OpenRouter API
  • What people are asking
  • Bottom line
  • Related on explainx.ai
← Back to blog

explainx / blog

OpenRouter Ox Alpha: Free 1M-Context Stealth Model for Coding Agents

OpenRouter, Stealth Models, AI Coding, Agent Harnesses, Model Routing, API

OpenRouter shipped Ox Alpha (stealth/ox-alpha) on Aug 20 — a free 1M-context stealth model for coding and agentic work. Specs, privacy caveats, and API setup.

Aug 21, 2026·12 min read·Yash Thakker
add explainx.ai
go deep
OpenRouter Ox Alpha: Free 1M-Context Stealth Model for Coding Agents

Update — September 5, 2026: OpenCode shipped a new stealth model, Omen Alpha, paid-only via OpenCode Go at ~$0.20/M input tokens — same anonymous-launch pattern, unconfirmed identity. Full coverage.

Update — August 26, 2026 (evening): Official name is GLM-5.3-Flash — 320B-A18B, MIT, $0.15/$0.50 API. Full launch guide.

Update — August 26, 2026: Z.AI (Zhipu) confirmed to Bloomberg that Ox Alpha is a new GLM-series model and said open weights drop tonight. Ox Alpha hit #1 on OpenRouter, more than doubling DeepSeek usage — community GLM-5.3-Flash theory vindicated. Full timeline: Ox Alpha identity post.

Update — August 25, 2026: By the end of Ox Alpha's first four-day free preview window, community dashboards and trade reporting cite up to ~26 trillion tokens processed on OpenRouter — roughly 2.6× the prior record for a single-model launch (CryptoBriefing reported 11.6T in 72 hours mid-window). Treat headline totals as platform-analytics estimates, not audited revenue; the takeaway for builders is demand shape — Claude Code, Hermes Agent, and Zed drove agentic, high-context burns while the model was $0/$0. See also Ox Alpha top use cases.

OpenRouter added a new stealth model on August 20, 2026: Ox Alpha (stealth/ox-alpha) — positioned as a frontier-class model for efficient coding, sustained agentic work, and production workloads, with a 1M-token context window, multimodal input, tool calling, and $0/$0 pricing during preview.

The release lands one day after AT&T reported 56% coding-cost savings from model routing and in the same week Stripe's investor letter named OpenRouter its largest-ever acquisition — still expected to close in the coming weeks. Ox Alpha is not about that deal; it is a concrete model you can point an agent harness at today, for free, through the same OpenRouter API many teams already use for Fusion panels and cost-aware routing.

TL;DR — what builders are asking

table · 2 cols
QuestionDirect answer
Model ID?stealth/ox-alpha
Released?August 20, 2026 (per OpenRouter model page)
Price?Free — $0/M input, $0/M output
Context / max output?1,048,576 tokens in · 131,072 tokens out
Modalities?Text, images, and video in → text out
Tool calling?Yes — supported; OpenRouter reports ~4.45% tool-call error rate (3-day avg)
Throughput / latency?~50 tok/s and ~2.02s P50 latency (OpenRouter dashboard)
Trains on your data?OpenRouter says prompts/completions are retained but not used for training (this release); OpenCode separately advertises Zero Data Retention for its own Ox Alpha traffic — see below
Who made it?Z.AI → GLM-5.3-Flash — MIT, 320B-A18B (launch)
Top agent traffic?Claude Code (~9.3B tokens), Hermes Agent (~9.0B), plus Oh-My-Pi, DeepSeek Harness, Z Code on the model's apps chart
How does it benchmark?Independent test by developer Ben Davis: 80% on DeepSWE, ahead of Fable (65%) and GPT-5.6 Sol (52%) — not an official OpenRouter or provider number
Free anywhere else?Yes — OpenCode (the open-source coding agent) also routes to Ox Alpha directly, and OpenCode Go made it near-unlimited and free for 6 more days as of Aug 21 evening, not counted against normal Go usage
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

A quick-watch rundown of the Ox Alpha stealth model.

What OpenRouter stealth models are

OpenRouter's stealth models are preview releases from third-party providers who stay anonymous during the preview window. OpenRouter routes traffic; it does not develop or own the weights.

The pattern is established:

  • Models ship under codenames (Ox Alpha, past examples include Hunter Alpha, Healer Alpha, Owl Alpha).
  • They are often free during preview to drive eval traffic.
  • Providers historically logged prompts for improvement — Ox Alpha's page makes a narrower claim for this release (more below).
  • OpenRouter later reveals some identities — Hunter and Healer turned out to be Xiaomi MiMo models. Ox Alpha was confirmed as Zhipu GLM on August 26, 2026 (Bloomberg) — open weights promised the same night.

Governance falls under OpenRouter's Stealth Model Terms — separate from standard provider terms. Treat stealth previews as time-boxed experiments, not long-term infrastructure commitments.

Ox Alpha specs (verified from OpenRouter)

These numbers come from the Ox Alpha model page as of August 21, 2026:

table · 2 cols
SpecValue
Model slugstealth/ox-alpha
Provider labelStealth (single upstream host)
Listed price$0 / $0 per million tokens
Context window1,048,576 tokens
Max output131,072 tokens
Input modalitiesText, image, video
OutputText
ReasoningSupported (reasoning parameter documented)
Tool callingSupported
Release dateAugust 20, 2026
Cache hit rate (3d avg)~67%
Uptime (3d)99.99%
Availability (3d)96.79%

OpenRouter's description emphasizes long-horizon software engineering, complex reasoning, and workflows that combine text with visual context — language that matches how agent harnesses actually burn tokens: repo maps, screenshots, test logs, and multi-step tool loops.

Who is already using it

OpenRouter publishes app-level token share for each model. On Ox Alpha's first days, the leaderboard is dominated by agent harnesses, not chat UIs:

  1. Claude Code — Anthropic's agentic coding tool (~9.32B tokens on the model chart)
  2. Hermes Agent — Nous Research's persistent open-source agent (~8.98B tokens)
  3. Oh-My-Pi, DeepSeek Harness, and Z Code — also listed among top senders on the model page

That traffic mix is a signal: teams running real coding agents are treating Ox Alpha as a workhorse model, not a playground curiosity. It does not prove quality on your repo — only that production-shaped workloads are flowing.

An independent benchmark: 80% on DeepSWE

Developer Ben Davis ran Ox Alpha through DeepSWE, a coding-agent benchmark, and reported it scoring 80% — ahead of Fable at 65% and GPT-5.6 Sol at 52%. This is not an OpenRouter or provider-published number; treat it the same way you'd treat any single developer's eval run: a useful data point, not a verified leaderboard entry. It does line up with the traffic pattern above — a model agent harnesses are routing real work to tends to also do well on agent-shaped benchmarks like DeepSWE.

Free pricing and the "no training" claim

Two details matter for anyone routing proprietary code through Ox Alpha.

Free — for now

OpenRouter lists Ox Alpha at $0/$0. Stealth previews have historically been free to accumulate feedback, then either graduate to a named model with list pricing or disappear. Budget for the possibility that "free" is a preview subsidy, not a permanent tier — the same caution AI token pricing guides apply: watch usage dashboards and keep fallback routes.

Retention ≠ training (read both lines)

OpenRouter's Ox Alpha banner states:

Prompts and completions for this model are retained by the provider and are not used for training; all other use is governed by the Stealth Model Terms.

That is a meaningful distinction:

table · 3 cols
ClaimWhat it impliesWhat it does not imply
Not used for trainingProvider says your completions won't fine-tune weights (this release)Prompts aren't stored
Retained by the providerLogs likely persist on provider infrastructureYou know who the provider is
Stealth Model TermsOther uses (eval, abuse monitoring, legal holds) may still applyEnterprise ZDR / HIPAA

For production secrets, an anonymous provider with retained logs is a different risk profile than calling Anthropic or OpenAI under a signed enterprise agreement — even when the sticker price is zero. The Databricks cost post and AT&T routing story both assume internal evals and tiered routing precisely because model substitution is never just a price question.

OpenCode's version: Zero Data Retention, and a "100T tokens/day" capacity claim

Ox Alpha isn't only accessible through OpenRouter. OpenCode — the open-source coding agent whose traffic already shows up among Ox Alpha's top senders above — announced its own direct route to the model on X:

"Ox Alpha (stealth model) is free for the next week — 1M Context, Multi-modal, Zero Data Retention. Generous rate limits, near unlimited usage. We have capacity for 100T tokens per day, let's see what you can do."

Two things are worth flagging rather than repeating at face value:

  • "Zero Data Retention" vs. OpenRouter's "retained but not used for training." These are different claims from different parties in the same request chain. OpenRouter's model page describes the upstream provider's retention policy; OpenCode's ZDR claim describes what OpenCode itself does with traffic that flows through its client. Neither statement contradicts the other outright, but they are not interchangeable — a ZDR promise from the client layer doesn't override a retention policy at the model-provider layer underneath it. If retention is a hard requirement for you, verify the full path, not just the layer that says the reassuring thing.
  • The 100 trillion tokens/day figure drew immediate skepticism. Theo (t3.gg) reacted publicly: "'We have capacity for 100T tokens per day' — okay who the fuck made this model and where did they get this much compute?" That's a fair question. For scale, 100T tokens/day is in the range of what only the very largest frontier-model providers claim to serve in total, across every model and customer combined — a single free stealth preview claiming that ceiling is either an extraordinary infrastructure story or marketing rounding. Nobody involved has clarified which.

OpenCode Go — OpenCode's hosted/managed offering — extended the same deal on August 21: near-unlimited, completely free access to Ox Alpha for 6 more days, and usage through it does not count against your normal OpenCode Go quota. If you already run OpenCode Go, this is currently the cheapest way to put real repo work through Ox Alpha without touching your regular usage budget.

Who is Ox Alpha? (GLM-5.3-Flash — August 26, 2026)

Official product name: GLM-5.3-Flash. Z.ai's @Zai_org launch post (Aug 26, 2026, 7:42 PM) states it was "previously previewed as Ox Alpha" on OpenCode and OpenRouter.

table · 2 cols
SpecValue
Architecture320B total / 18B active
Context1M tokens, natively multimodal
LicenseMIT — Hugging Face
API price$0.15/M in · $0.50/M out · $0.03/M cached
Stealth infraEntire preview on Chinese AI chips, per Z.ai
Local inferenceSGLang, vLLM, TokenSpeed

Full benchmarks, architecture, and routing guide: GLM-5.3-Flash launch post. Forensics timeline: Ox Alpha identity post.

How to try Ox Alpha via the OpenRouter API

OpenRouter's API is OpenAI-compatible. Swap the base URL and model slug.

1. API key

bash
export OPENROUTER_API_KEY="sk-or-v1-..."

Create keys in the OpenRouter dashboard.

2. Minimal curl (non-streaming)

bash
curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "stealth/ox-alpha",
    "messages": [
      {"role": "user", "content": "Refactor this function to handle nil inputs safely."}
    ]
  }'

3. Agent harness / Claude Code / Hermes

Point your harness provider at OpenRouter and set the model to stealth/ox-alpha. Optional headers (HTTP-Referer, X-Title) only affect OpenRouter leaderboard attribution — not model behavior.

For tool-heavy loops, enable streaming and watch cumulative latency: OpenRouter reports ~11.6s median end-to-end latency on multi-step agent turns — acceptable for background agents, painful for tight interactive edits unless you route subtasks to faster tiers.

4. Routing pattern (recommended)

Mirror enterprise playbooks from AT&T's LiteLLM gateway:

text
cheap/open route (Ox Alpha) → internal eval gate → premium model on failure

Ox Alpha is a strong candidate for the cheap tier while it stays free — not a blind default for merges, security reviews, or regulated data.

What people are asking

Is Ox Alpha actually "frontier"?

OpenRouter's marketing calls it frontier-class for efficient coding and agentic work. There is no independent benchmark sheet on the model page — only usage stats and latency. Community threads compare it favorably to some Chinese open models and note fast tool use, but your golden tasks (lint fixes, migrations, test generation) are the only eval that matters.

Will it stay free after Stripe closes the OpenRouter deal?

Unknown. Stripe's acquisition is not closed yet; Ox Alpha pricing is set by the anonymous provider + OpenRouter listing, not by Stripe today. Watch the model page and your invoice — free stealth previews can change overnight.

Ox Alpha vs OpenRouter Fusion

table · 3 cols
Ox AlphaOpenRouter Fusion
ShapeSingle stealth modelMulti-model panel + judge
CostFree (preview)Sum of panel + judge tokens
Best forHigh-volume agent loops, codingResearch, deliberation, high-stakes synthesis
IdentityAnonymousNamed panel models

They solve different problems. Fusion is for when being wrong is expensive; Ox Alpha is for when token volume is expensive — at least while the preview lasts.

Should I send customer PII or prod secrets?

No — not to an anonymous stealth provider that retains logs. Use Ox Alpha for sanitized repos, open-source work, or spikes where log retention is acceptable. Same rule as any free preview tier.

Bottom line

Ox Alpha is real, free, and already carrying billions of tokens from Claude Code and Hermes Agent — and Zhipu confirmed on August 26, 2026 that it is a GLM model with open weights releasing that night. The actionable move: add stealth/ox-alpha as a routed tier (or migrate to the named Z.AI slug once OpenRouter updates), run your coding evals, and watch for the Hugging Face drop if you self-host.

Related on explainx.ai

  • Omen Alpha: OpenCode's new stealth model at $0.20/M tokens — the sequel stealth launch, September 2026
  • Top 10 things people are building with Ox Alpha — fluid sims, 3D scenes, a DeepSWE benchmark run, and more
  • GLM-5.3-Flash official launch — Ox Alpha unmasked
  • Ox Alpha: forensics timeline
  • Why free preview pricing wrecks model evaluation — the corrected Gemini/Ox Alpha timing timeline, plus a GLM-5.3-Flash vs Gemini 3.7 Flash decision table
  • Stripe acquires OpenRouter — what builders should know
  • AT&T cut AI coding costs 56% with model routers
  • Databricks: managing AI coding costs at scale
  • OpenRouter Fusion API guide
  • Hermes Agent #1 on OpenRouter rankings
  • AI token pricing, explained
  • Choosing open-weight vs closed models
  • What is OpenRouter? Enterprise guide · Loop engineering with coding agents
  • Inkling's "free on OpenRouter" headline, decoded — the fine print

Primary sources: OpenRouter — Ox Alpha model page · OpenRouter — Stealth provider · OpenRouter — Stealth Model Terms


Accurate as of August 26, 2026. Zhipu confirmed GLM lineage to Bloomberg; open weights promised Aug 26 night. Stealth slug, pricing, and limits may change after reveal — see identity timeline.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Sep 5, 2026

Omen Alpha: OpenCode's New Stealth Model at $0.20 Per Million Tokens

OpenCode added a second anonymous stealth model, Omen Alpha, to OpenCode Go on September 4, 2026 — a 500K-context coding model with reference pricing of $0.20 per million input tokens. This is the same playbook OpenCode ran with Ox Alpha in August, right down to the unconfirmed lab speculation.

Aug 21, 2026

OpenRouter vs. Direct Provider APIs: An Enterprise Decision Guide (2026)

OpenRouter's pitch is one integration for 400+ models with automatic fallback and unified billing; going direct to a provider gets you the lowest latency, no middleman fee, and day-one access to new features. Here's the actual decision framework — team size, model count, cost vs. reliability, compliance — plus what changed after Stripe's $7B acquisition.

Aug 21, 2026

Top 10 Things People Are Building With Ox Alpha

OpenRouter's free, anonymous Ox Alpha model has been live for less than 48 hours, and builders are already testing it against everything from GPU-accelerated physics to full desktop-environment clones. Here are the ten most notable things people have actually built and shared — plus the honest caveats where the "wow" factor doesn't hold up to scrutiny.