SDSignal Desk

StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agentic Work

Sep 20, 2026, 11:37 PM · MarkTechPost

Image: MarkTechPost

StepFun’s Step 5 Preview bets on a 600B MoE with 27B active and a 1M context — selling agentic work on price, with open weights promised mid-October.

Why it matters

Another frontier-class sparse model is hitting API first: Step 5 Preview targets software engineering, knowledge work, and finance with a cost pitch. Hosted access is live; open weights are scheduled for October 15, 2026.

For teams building long-horizon agents, a million-token context plus aggressive cache pricing can change what “keep the whole project in prompt” costs — if quality holds outside vendor benches.

From the desk

We’re seeing the familiar MoE bargain: huge total capacity, thin active path per token. StepFun lists about 600B total parameters and 27B active — roughly 4.5% per token — with text, image, and video in, text out, and reasoning effort controls. Streaming, tool calling, JSON mode, and prompt caching are on the feature list for agent builders.

Company-reported scores put it near but usually behind Claude Opus 5 and GPT-6 Astra on several agent and coding boards. Artificial Analysis scores it 44 on its Intelligence Index versus a 24 median in a similar price tier, with about 99.8 tokens per second on the API. Pricing is the sharp edge: $1.00 per million input tokens ($0.05 cache hit) and $2.70 output, well under medians Artificial Analysis cites for peers — though verbose reasoning can claw back savings when output tokens balloon.

Useful AI at lower task cost is a real public good if independent evals keep confirming it. Vendor benches and a house benchmark like StepCodeBench deserve skepticism until outsiders reproduce them. The architecture story — narrow-deep 92-layer stack, long-horizon RL, speculative decoding — is plausible engineering color, not a reason to skip bake-offs.

Our take: watch the October 15 open-weight date. API-only “preview” launches are easy; shipping weights that others can serve is the credibility test. Self-hosters should plan for multi-GPU footprints once weights land. I’m also watching whether 1M context is usable in practice or mostly a marketing ceiling once KV-cache and latency show up.

If this pattern becomes normal, agentic workloads keep migrating to whichever MoE offers enough intelligence per dollar — and closed labs feel price pressure even when they lead on raw scores.

Context

MarkTechPost report by Michal Sutter, Sep 20, 2026, summarizing StepFun’s Step 5 Preview specs, pricing, reported benchmarks, and Artificial Analysis measurements.

Who feels it

Agent builders
Cheaper long-context API option for multi-step coding and research agents; verify quality on internal evals.
Enterprises in finance and knowledge work
Cost-shaped alternative to top closed models if compliance and reliability checks pass.
Self-hosting teams
Plan multi-GPU footprints for ~600B weights after the scheduled October open release.
Rival labs
Price/performance pressure on API tiers marketed for agentic workloads.

What to watch

  1. Whether open weights actually ship on October 15, 2026.
  2. Third-party reproductions of FrontierFinance, DRACO, and coding scores.
  3. Real-world cost after reasoning-token verbosity on production agent traces.
  4. Quality of the Claude Code / Step Plan integration in developer hands.

Read the original

Continue at the source.

MarkTechPost