StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agentic Work
Sep 20, 2026, 11:37 PM · MarkTechPost
StepFun’s Step 5 Preview bets on a 600B MoE with 27B active and a 1M context — selling agentic work on price, with open weights promised mid-October.
Why it matters
Another frontier-class sparse model is hitting API first: Step 5 Preview targets software engineering, knowledge work, and finance with a cost pitch. Hosted access is live; open weights are scheduled for October 15, 2026.
For teams building long-horizon agents, a million-token context plus aggressive cache pricing can change what “keep the whole project in prompt” costs — if quality holds outside vendor benches.
From the desk
We’re seeing the familiar MoE bargain: huge total capacity, thin active path per token. StepFun lists about 600B total parameters and 27B active — roughly 4.5% per token — with text, image, and video in, text out, and reasoning effort controls. Streaming, tool calling, JSON mode, and prompt caching are on the feature list for agent builders.
Company-reported scores put it near but usually behind Claude Opus 5 and GPT-6 Astra on several agent and coding boards. Artificial Analysis scores it 44 on its Intelligence Index versus a 24 median in a similar price tier, with about 99.8 tokens per second on the API. Pricing is the sharp edge: $1.00 per million input tokens ($0.05 cache hit) and $2.70 output, well under medians Artificial Analysis cites for peers — though verbose reasoning can claw back savings when output tokens balloon.
Useful AI at lower task cost is a real public good if independent evals keep confirming it. Vendor benches and a house benchmark like StepCodeBench deserve skepticism until outsiders reproduce them. The architecture story — narrow-deep 92-layer stack, long-horizon RL, speculative decoding — is plausible engineering color, not a reason to skip bake-offs.
Our take: watch the October 15 open-weight date. API-only “preview” launches are easy; shipping weights that others can serve is the credibility test. Self-hosters should plan for multi-GPU footprints once weights land. I’m also watching whether 1M context is usable in practice or mostly a marketing ceiling once KV-cache and latency show up.
If this pattern becomes normal, agentic workloads keep migrating to whichever MoE offers enough intelligence per dollar — and closed labs feel price pressure even when they lead on raw scores.
Context
MarkTechPost report by Michal Sutter, Sep 20, 2026, summarizing StepFun’s Step 5 Preview specs, pricing, reported benchmarks, and Artificial Analysis measurements.
Who feels it
- Agent builders
- Cheaper long-context API option for multi-step coding and research agents; verify quality on internal evals.
- Enterprises in finance and knowledge work
- Cost-shaped alternative to top closed models if compliance and reliability checks pass.
- Self-hosting teams
- Plan multi-GPU footprints for ~600B weights after the scheduled October open release.
- Rival labs
- Price/performance pressure on API tiers marketed for agentic workloads.
What to watch
- Whether open weights actually ship on October 15, 2026.
- Third-party reproductions of FrontierFinance, DRACO, and coding scores.
- Real-world cost after reasoning-token verbosity on production agent traces.
- Quality of the Claude Code / Step Plan integration in developer hands.