explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • The Pareto frontier claim, and what it actually shows
  • The second chart: where Step 5 Preview actually sits among named competitors
  • Architecture: 600B total, 27B active, and what that buys
  • Pricing and availability today
  • The same week: classifier.dev and JEPA-Anything
  • Honest limitations
  • What this means for builders
  • Related on explainx.ai
← Back to blog

explainx / blog

StepFun Step 5 Preview: A New Pareto Frontier for Agentic AI

StepFun, Model Launches, Mixture of Experts, Artificial Analysis, Open Source AI, AI Benchmarks

StepFun's Step 5 Preview claims a new cost/intelligence Pareto frontier — 600B/27B MoE, 1M context, vision, open weights October 15, 2026.

Sep 20, 2026·11 min read·Yash Thakker
add explainx.ai
go deep
StepFun Step 5 Preview: A New Pareto Frontier for Agentic AI

Model launches in September 2026 keep making the same claim — cheaper, not just smarter. StepFun's Step 5 Preview, announced September 20, 2026, is the latest to lean on that framing hard: not "we beat GPT-6 Astra," but "we moved the cost/intelligence Pareto frontier itself." That's a more precise and more checkable claim than most launch-day superlatives, and StepFun backed it with two Artificial Analysis charts rather than a bare percentage. Here's what the charts actually show, what's confirmed versus what's still preview-only, and where Step 5 Preview lands next to GLM-5.3, Kimi K3, and DeepSeek V4.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR

table · 2 cols
QuestionAnswer
What is it?StepFun's new flagship model for agentic software engineering and finance-heavy knowledge work
Architecture600B total parameters, ~27B active (Mixture-of-Experts)
Context window1 million tokens, with vision support
Headline claimA new cost/intelligence Pareto frontier on Artificial Analysis — roughly 65% lower cost at matched intelligence versus the prior efficient-tier line
Does it beat GPT-6 Astra or Claude Opus 5 outright?No — those remain higher on raw Intelligence Index score, at much higher cost per task
Open weights?Not yet — StepFun states October 15, 2026
How to try it nowHosted API at platform.stepfun.ai
Model pagestepfun.com/step-5-preview

The Pareto frontier claim, and what it actually shows

StepFun's launch tagline is "Advancing the Pareto Frontier" — a specific, testable framing borrowed straight from the Artificial Analysis methodology that explainx.ai has covered before. A Pareto frontier, in this context, is the curve connecting the best available Intelligence Index score at each price point — a model only earns a spot on the frontier if no cheaper model matches its score and no pricier model at the same score beats it. Claiming to move that line is a claim about efficiency, not about topping the leaderboard outright.

StepFun Step 5 Preview cost/intelligence Pareto frontier chart showing the new frontier line pulling left of the previous curve at roughly 65% lower cost

Chart: StepFun's own published Pareto frontier comparison, via stepfun.com/step-5-preview.

The chart StepFun published plots the Artificial Analysis Intelligence Index against cost per task (USD, log scale), split into three informal bands — efficient models, workhorse models, and high-compute models. Step 5 Preview sits right at the boundary between the efficient and workhorse tiers, marked as pulling the frontier line up and to the left of where DeepSeek V4 Pro (0813, max) previously anchored the "previous Pareto line." StepFun's annotation puts the delta at -65% cost for a comparable point on the new curve versus the old one. GPT-6 Astra (max) and Claude Opus 5 (max) remain further right on the chart — higher intelligence score, but at multiples of the cost — meaning Step 5 Preview isn't claiming to unseat the highest-compute frontier models, only to make a specific mid-tier cost/intelligence tradeoff better than it previously was.

The second chart: where Step 5 Preview actually sits among named competitors

StepFun's second published chart is more useful for a direct competitive read, because it names specific rival models rather than showing an abstracted frontier line.

Intelligence Index vs cost per task scatter chart naming StepFun Step 5 Preview alongside GLM-5.3-Flash, DeepSeek V4.1 Flash, Kimi K3, GPT-6 Astra, and Claude Opus 5

Chart: StepFun's published Intelligence Index vs. cost-per-task scatter, via stepfun.com/step-5-preview.

Reading the chart's own dot placements:

table · 3 cols
ClusterModels in itWhat it means for Step 5 Preview
"Most attractive quadrant" (StepFun's own shaded region)Step 5 Preview, GLM-5.3-Flash, DeepSeek V4.1 Flash, GPT-5.6-Luna (max)Step 5 Preview sits just above GLM-5.3-Flash and DeepSeek V4.1 Flash on the Intelligence Index axis, at a similar cost-per-task band — the direct "efficient tier" comparison set
Mid-cost clusterKimi K3 (max), GPT-6 Astra (max), GLM-5.3 (max), Qwen3.8 Max (0902), Grok 4.6 (high)Materially higher cost-per-task than Step 5 Preview, with higher Intelligence Index scores — Step 5 Preview does not claim to beat this tier
High-compute clusterClaude Opus 5 (max), Claude Fable 5.1 (max with fallback)The most expensive, highest-scoring cluster on the chart — outside Step 5 Preview's claimed comparison set entirely
Below the frontierMistral Medium 3.5, MiniMax-M3, gpt-oss-120b (high), Gemini 3.5 Flash-Lite, Muse Glimmer (high)Lower Intelligence Index at similar or higher cost than Step 5 Preview — StepFun's chart positions these as dominated points

The honest read: Step 5 Preview's claim is specifically an efficient-tier claim, contested directly against GLM-5.3-Flash and DeepSeek V4.1 Flash, not against the mid-cost or high-compute clusters where GLM-5.3 (max), Kimi K3 (max), GPT-6 Astra, and Claude Opus 5 sit. That's a narrower and more defensible claim than "we beat everything," and it lines up with how explainx.ai's benchmark-literacy guide recommends reading any vendor-published chart — check which comparison set the chart actually draws before accepting the headline framing.

Architecture: 600B total, 27B active, and what that buys

Step 5 Preview is a Mixture-of-Experts (MoE) model with 600 billion total parameters and roughly 27 billion active per token — a sparsity ratio (about 4.5% of parameters active per forward pass) that's in the same general family as other 2026-era MoE flagships, including GLM-5.3's 744B/40B design and DeepSeek V4's dense-vs-sparse tradeoffs. The MoE approach lets StepFun claim frontier-adjacent capability while keeping per-token compute — and therefore per-task cost — closer to a much smaller dense model, which is the architectural mechanism underneath the "lower cost per task" half of the Pareto frontier claim.

Two additional specs matter for how the model actually gets used:

  • 1M-token context window — enough for large codebases, long agent transcripts, or multi-document financial analysis in a single call, without external retrieval scaffolding for many practical workloads.
  • Vision support — Step 5 Preview accepts image input alongside text, positioning it for multimodal agentic tasks (reading a screenshot, a chart, or a scanned document as part of an agent loop) rather than text-only pipelines.

StepFun's own description frames the model around sustained execution over long horizons — an agent-specific capability claim distinct from raw benchmark score, meaning the model is tuned to keep track of prior results and tool outputs across a long-running task rather than losing coherence as context accumulates. StepFun states Step 5 Preview coordinated 950 web fetches in a single agent action during one of its own demonstrated research workflows — a concrete, specific number, though self-reported and not yet independently reproduced.

Pricing and availability today

Step 5 Preview is live now, API-only, at platform.stepfun.ai — StepFun hasn't published a standalone per-token price sheet for Step 5 Preview independent of the cost-per-task figures shown in its own Artificial Analysis-style charts, so builders evaluating it today should benchmark against their own representative workload rather than the chart's aggregate cost-per-task number alone, the same caution explainx.ai gives for any newly launched model.

Open weights are promised for October 15, 2026 — about a month out from this preview launch. Until that date, Step 5 Preview is not self-hostable, and none of its architecture claims (the exact MoE routing design, training data composition, or Artificial Analysis chart methodology) can be independently verified by outside researchers. That's a meaningfully different position than GLM-5.3 or DeepSeek V4, both of which already ship downloadable weights — Step 5 Preview's efficiency claim currently rests entirely on StepFun's own hosted-API numbers and self-published charts.

The same week: classifier.dev and JEPA-Anything

Step 5 Preview wasn't the only fast-moving release this week. Two other stories are worth a mention, even though neither gets the same depth here:

classifier.dev launched as a free, zero-shot text classification API — no API key, no account, plain HTTP. Its own published benchmark table claims it outperforms TypeSafe AI's Jev on public AG News and emotion classification subsets: 90.0% versus Jev's 87.5% aggregate on AG News, with a wider gap (83.7% vs. 65.3%) on the harder "unsure" subset where Jev's own confidence calibration flags uncertainty. The service runs on Jev itself as its underlying model — its pitch is a "smart tier" that re-asks only the items Jev was unsure about, rather than a competing base model. That's a genuinely interesting data point for explainx.ai's ongoing Jev vs. XGBoost and BERT comparison: a wrapper around Jev's own uncertainty signal, not a from-scratch competitor, closing part of the accuracy gap that comparison identified. It's also fully open source, per its own launch thread, with the creator explicitly warning the service was freshly deployed and might break — worth checking current uptime before relying on it in production, and treating the benchmark numbers as self-reported and freshly published rather than independently reproduced.

JEPA-Anything, published on alphaXiv, proposes a single world-modeling framework spanning vision, biology, control, molecules, physics, weather, and clinical data — a considerably broader scope than the video-specific JEPA work explainx.ai covered with LeVJEPA's 20x compute-efficiency result. Its core technique, Orthogonal Predictive Factorization, splits a single JEPA latent state into complementary factors that each predict a different part of the world and recombine into a full state — improving results across all 10 matched dynamics tasks the paper tests, and reportedly recovering Kepler's scaling law from the raw latent factors along with nominating a biologically validated intervention. Both claims are striking enough to warrant independent scrutiny before treating them as settled; explainx.ai will revisit JEPA-Anything in more depth if the code repository referenced in the paper's discussion becomes public.

Honest limitations

  • All benchmark and pricing figures in this post are StepFun's own published charts — no independent, third-party reproduction of the Pareto frontier or Intelligence Index vs. cost claims had occurred as of this post.
  • Open weights are not yet available, so Step 5 Preview's architecture claims (exact MoE routing, training data composition) can't currently be verified outside StepFun's own disclosures.
  • The "950 web fetches in one agent action" figure is self-reported in StepFun's own launch thread, without a published methodology for what counted as a single "agent action."
  • classifier.dev's benchmark numbers are freshly self-published, from a service its own creator described as newly deployed and not yet stress-tested at scale.
  • JEPA-Anything's Kepler's-law and biological-intervention claims come from the paper itself, not an independent replication — treat them as a striking preliminary result, not a settled finding.

What this means for builders

If you're choosing a model for agentic software engineering or finance-heavy workloads today, Step 5 Preview is worth evaluating specifically if your workload sits in the efficient-tier cost band — where it's directly competing with GLM-5.3-Flash and DeepSeek V4.1 Flash on StepFun's own chart — rather than as a drop-in replacement for GPT-6 Astra, Claude Opus 5, or Kimi K3 at the high-compute end. Run your own representative task set against StepFun's hosted API before committing, since the published cost-per-task and Intelligence Index figures are aggregate numbers from a single vendor's chart, not a controlled, task-matched benchmark against your specific workload. If self-hosting matters to your deployment, the October 15, 2026 open-weights date is the milestone to watch — until then, GLM-5.3, DeepSeek V4, and Kimi K3 remain the downloadable alternatives in the same general cost tier.

Related on explainx.ai

  • Artificial Analysis Intelligence Index v4.2: What Actually Changed — the methodology behind the Index StepFun's charts are built on
  • GLM-5.3's "50% Coding Boost" Explained — the closest open-weight competitor in Step 5 Preview's efficient-tier comparison set
  • How to Read AI Benchmarks Without Getting Fooled — the framework this post applies to StepFun's own charts
  • Jev vs. XGBoost and BERT Classifiers: Is a System One Model Actually New? — where classifier.dev's Jev-beating claim fits
  • LeVJEPA: Video Pretraining at 20x Less Compute — the JEPA research lineage JEPA-Anything extends
  • What Are World Models? A Complete Guide — background on the world-modeling category JEPA-Anything targets
  • OpenJev: Try Jev's Trick Yourself, Free, in Your Browser — another free, Jev-adjacent tool from the same fast-moving week

Primary sources: StepFun Step 5 Preview model page · StepFun platform · Artificial Analysis · classifier.dev · JEPA-Anything on alphaXiv


Benchmark figures, pricing, and the October 15, 2026 open-weights date reflect StepFun's own published materials as of this post's September 21, 2026 update. No independent, third-party reproduction of StepFun's Pareto frontier or Intelligence Index claims had occurred as of publication — verify against your own representative workload before making a production decision.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 1, 2026

GLM-5.3's "50% Coding Boost" Explained — What Z.ai Actually Measured

Z.ai's headline claim that GLM-5.3 is "50% better at coding" than GLM-5.2 traces to one specific benchmark — Z.ai's own in-house Code Bench, at its highest reasoning tier — not a blanket coding-performance jump. The gains are real and the model reuses GLM-5.2's exact base weights, but the license also quietly changed in a way self-hosters should read closely.

Sep 21, 2026

Mozilla: Open-Weight Models Now 4 Months Behind Frontier

A Sept. 20, 2026 AI news digest surfaced a bare headline: Mozilla says the gap between the best open-weight model and the best closed frontier model has narrowed to about 4 months. There is no linked report to verify it against yet — so this post explains what a "gap in months" claim actually needs to measure to be credible, using explainx.ai's own benchmark coverage as the grounding.

Aug 27, 2026

Gemini's Ox Alpha Timing Backlash: The Corrected Timeline

Three posts from Google Gemini team members got read as trolling a rival model's launch. The dates say otherwise: the posts are from August 22, 2026, and Z.ai did not reveal Ox Alpha as GLM-5.3-Flash until August 26. Logan Kilpatrick's public reply was substantially correct — and underneath the drama sits a real practitioner question about how free preview windows distort model evaluation.