Prime Intellect opened Prime Inference on October 2, 2026 — the production serving stack it already used for reinforcement-learning rollouts, synthetic data, evaluations, and long-running coding agents. The product is not another agent harness. It is the missing serving layer in the stack explainx.ai has already covered as Prime Agent, open-source RL post-training, and Prime Sandboxes.
Primary source: Prime Inference: Fast, Reliable Serving for Frontier Open Models. Docs: Prime Inference overview.
Feed blurbs that say the company "served trillions of tokens" are pointing at Prime's own claim of nearly a trillion tokens per day internally before the public launch — not a separate, unverified public-traffic figure. Correct for that when you quote the headline.
TL;DR — what people are asking
| Question | Answer |
|---|---|
| What shipped? | Public Prime Inference: serverless + reserved capacity for frontier open models |
| When? | October 2, 2026 (announcement); GLM-5.3 public on OpenRouter since September 22 |
| API? | OpenAI-compatible at https://api.pinference.ai/api/v1 |
| First hosted model? | z-ai/glm-5.3 on Prime's own GPU fleet (NVIDIA Blackwell today) |
| Pricing public? | Token billing; per-model rates in catalog/API. Aggregators (~Oct 2): ~$1.40 / $4.40 per 1M in/out for GLM-5.3 |
| Vs Prime Agent? | Agent = harness. Inference = tokens under SLA |
| Vs prime-rl? | Trainer/post-training. Inference = serve the policy |
| Vs Sandboxes? | Isolated exec for tools/rollouts. Inference = the model endpoint |
| Use this week? | Swap base_url, run agent/eval traffic on GLM-5.3, measure tool-call errors + TTFT |
What Prime Inference actually is
Prime's mission framing in the launch post is blunt: post-training is incomplete until a trained model can serve real users, generate new experience, and feed production traces back into training. Prime Inference is that serve step.
Concretely, the product is:
- Hosted models — Prime runs the weights on its own inference infrastructure. GLM-5.3 is the first public hosted example.
- Gateway models — same API key and endpoint, but the request is routed to an external provider (docs make this distinction explicit).
- Two capacity shapes — serverless for variable demand; reserved capacity for sustained workloads.
- Ops features that open-weight self-hosting usually skips — production SLAs, security/privacy defaults, unified billing, team usage tracking, automatic failover across datacenters.
The stack underneath is named in the announcement: NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer, developed with Inferact and NVIDIA, with selected improvements contributed upstream. That is a production serving story, not a research notebook.
If you already think in terms of loop engineering — generate, evaluate, fix, repeat — Inference is the generate hop that has to stay up while the rest of the loop burns tokens.
How to call it (minimal path)
Official docs walk a three-step path:
- Create an API key with the Inference permission (without it, requests fail auth).
- Export
PRIME_API_KEY. - Hit chat completions with a catalog model ID.
CLI:
# uv tool install prime && prime login
prime inference chat 'z-ai/glm-5.3' "Write a haiku about KV caches."
OpenAI-compatible HTTP:
export PRIME_API_KEY="your-api-key"
curl https://api.pinference.ai/api/v1/chat/completions \
-H "Authorization: Bearer $PRIME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "z-ai/glm-5.3",
"messages": [{"role": "user", "content": "Hello"}]
}'
Python SDK pattern from docs:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["PRIME_API_KEY"],
base_url="https://api.pinference.ai/api/v1",
)
response = client.chat.completions.create(
model="z-ai/glm-5.3",
messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)
Team billing defaults to the key owner's personal balance unless you send X-Prime-Team-ID. CLI team config does not automatically apply to raw HTTP/SDK calls — you must set the header (or default_headers on the client). List IDs with prime teams list or the team profile page.
List the live catalog with GET /models or prime inference models. Do not hard-code yesterday's model list into production config without a refresh path.
Pricing: what is public today
Prime bills by token against the account balance. Docs say pricing varies by model and should be read from the model catalog / Models API before you choose a serving type — hosted vs gateway rates and limits differ.
The October 2 launch post does not publish a fixed dollar table for every model. That is intentional: rates move with the catalog. Third-party trackers updated the same day listed GLM-5.3 on Prime Intellect at roughly $1.40 per 1M input, $4.40 per 1M output, and about $0.26 per 1M cached input. Treat those as aggregator snapshots, not a contract — re-check the catalog before you sign a budget or compare against another host.
Roadmap items that matter for cost later: batch and async inference for cheaper offline jobs, and one-click / reserved deployments of fine-tuned models from Prime training runs onto dedicated capacity. Neither is the mainline product on day one; both tell you where the pricing surface is headed.
For context on why reserved GPU economics keep moving under everyone, see explainx.ai's coverage of AWS Capacity Blocks rate changes. Inference APIs abstract the GPU invoice; they do not erase it.
Models and hardware claims (verified from the launch)
| Fact | Source claim |
|---|---|
| First public hosted model | GLM-5.3 (z-ai/glm-5.3) |
| OpenRouter soft launch | September 22, 2026 |
| Uptime claim since that soft launch | 100% (vendor-reported) |
| Tool-call quality claim | Near-zero error rate on their suite |
| Hardware today | NVIDIA Blackwell; Vera Rubin "coming soon" |
| Internal traffic before public launch | Nearly a trillion tokens/day |
| Customer prod serving | Since January 2026 |
| Benchmark target | ~100 end-to-end tok/s/user |
| Published GLM-5.3 GB200 NVL72 point | 1:4 prefill/decode → 66 sessions/prefill group at 101 tok/s/user |
Those numbers are vendor-run and model-specific. Prime says so. Re-measure on your prompt mix before you put the figures in a board slide.
Why this stack exists: agents break naive serving
The launch post's workload sketch matches what builders already feel:
- A typical agent turn adds about 6K tokens onto a ~140K-token prompt.
- Most of the history is reusable (KV cache), but cold long prompts arrive in the same fleet.
- Prefill (prompt processing) and decode (token generation) fight for the same GPUs if you do not split them.
- Silent tool-call failures are worse than slow tokens — the agent has nothing to execute and no error to retry.
That is the same failure mode space as long-horizon agent harness work: reliability of structured actions matters as much as raw tokens-per-second.
Prime's answer, summarized without turning this news post into a kernel paper:
- Prefill/decode disaggregation via Dynamo + vLLM + NIXL KV transfer — they report ~40% lower p90 inter-token latency in their tests vs shared-GPU chunked prefill.
- KV-aware routing plus Mooncake as a DRAM second cache tier so prefixes can live outside GPU HBM.
- NVFP4 compression on GLM-5.3's MLA cache rows (576 → 352 bytes) for ~50% more cached tokens at the same memory budget, with a native sparse-MLA kernel headed to FlashInfer.
- BLHNC block-major KV layout to cut NVLink transfer fragmentation (~10× fewer descriptors, ~47% lower mean transfer time in a cited TP8 comparison).
- Grammar-constrained tool calls — structural-tag builder for Dynamo/xgrammar so undeclared or malformed tool calls do not silently become a normal stop.
If you care about Apple-silicon local engines instead of hosted fleets, that is a different product category — see Perplexity Lily. Prime Inference is the cloud SLA side of the same "serving is the product" trend.
Prime Inference vs Prime Agent vs prime-rl vs Sandboxes
This is the section most readers need, because the brand shares a prefix and the products share a roadmap.
| Product | Job | You use it when… |
|---|---|---|
| Prime Inference | Production model serving (OpenAI-compatible) | You need tokens, SLAs, failover, team billing |
| Prime Agent | Open-source coding/research harness (RLM + Continual Harness) | You want an agent loop that edits its own supplemental state |
| prime-rl | Open-source RL trainer / post-training stack | You are updating weights from rollouts and rewards |
| Prime Sandboxes | Isolated execution for tools, evals, RL environments | You need safe concurrent code/browser/tool runs |
| Lab (adjacent) | Hosted training + evals that can deploy adapters | You want the managed loop that ends in a deployment |
Practical mapping:
- Point Prime Agent (or Claude Code, Pi, or your own harness) at Prime Inference as the model backend — harness ≠ endpoint.
- Run prime-rl / Lab training to produce adapters, then serve them through Inference (roadmap: one-click reserved deploy from training runs).
- Keep Sandboxes under the agent for tool execution; do not confuse sandbox concurrency with model TPS.
We covered the harness side in Prime Agent's RLM + Continual Harness launch and the trainer/environment side in open-source RL-as-a-service stacks (Miles, SkyRL, prime-rl). Sandbox scale showed up again in the DeepSeek DSec / Prime Sandboxes story. Today's news closes the serving gap those posts left open.
Pi remains the minimal harness ancestor in this neighborhood — useful if you want a thinner client against the same OpenAI-shaped API: Pi harness guide.
What a builder should do this week
Skip the kernel deep dive unless you are writing one. Ship a measurement:
- Create a key with Inference permission and confirm a tiny chat completion against
z-ai/glm-5.3. - Point your agent harness at
https://api.pinference.ai/api/v1without rewriting tools — onlybase_url+ key + model ID. - Run your own tool-call eval (required tools, nested schemas, streaming vs non-streaming). Prime's claim is near-zero error on their suite; yours is the only suite that matters.
- Capture TTFT and tok/s/user on a realistic long-context replay (returning 140K-class sessions plus cold arrivals), not a 20-token hello world.
- Decide capacity shape — serverless if traffic is spiky and you are still exploring; talk reserved if you already know you will burn steady GPU-hours for weeks.
- Wire team billing headers before finance asks why personal balances evaporated.
- If you train on Prime — watch the reserved/fine-tune deploy path on the roadmap; do not assume day-one parity with every Lab adapter.
Honest limitations from the public materials:
- Hosted catalog starts narrow (GLM-5.3 is the headline); gateway fills gaps but is not the same as "Prime runs it."
- Dollar prices are catalog-live, not a single immortal table in the blog post.
- Performance figures are vendor benchmarks on GB200 NVL72 with AgentX-style mixes.
- Batch/async and one-click fine-tune deploys are roadmap, not today's default button.
- You still need sandboxes, evals, and a harness — Inference does not replace them.
What people are asking
Is this just OpenRouter with extra steps?
No. Prime already listed GLM-5.3 on OpenRouter on September 22 as a distribution channel. Prime Inference is the first-party API, capacity model (serverless vs reserved), billing, and multi-DC failover story. You can use either surface; they are not the same product contract.
Can I self-host the same stack?
Pieces are open (vLLM, Dynamo, FlashInfer contributions). The product you are buying is the operated fleet, SLAs, and the agent-oriented optimizations already wired. Self-hosting those components is a different project with a different on-call bill.
Does this make closed-model APIs obsolete?
No. It makes open-weight production serving less of a DIY reliability tax. If your requirement is a specific closed model, you still need that vendor — or a gateway route if Prime lists it. If your requirement is open weights with tool-call SLAs, this is aimed at you.
Should I move RL rollouts onto the public API?
Prime built the stack for RL rollouts and still runs enormous internal volume. Public docs emphasize evaluations and chat completions. For heavy training rollouts, confirm with Prime whether your plan is serverless public, reserved, or still the internal/Lab path — do not assume the blog demo equals your training cluster.
Related on explainx.ai
- Prime Agent: RLM Continual Harness (Aug 2026) — the harness, not the inference API
- Open-source RL-as-a-service stacks — prime-rl next to Miles and SkyRL
- DeepSeek DSec / Prime Sandboxes — execution isolation at scale
- What is an agent harness?
- Loop engineering with coding agents
- Pi: minimal agent harness
- Perplexity Lily Apple Silicon inference — local engine contrast
- AWS Capacity Blocks GPU price hike — reserved GPU economics under the API
Primary sources: Prime Intellect announcement · Inference docs · Models API
Product facts, model IDs, capacity options, and performance numbers reflect Prime Intellect's October 2, 2026 announcement and docs as of October 3, 2026. Catalog pricing and available models change — verify against the live Models API before production use.
