Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogStoryContact

Step 5 Preview: Specs, Price, Benchmarks, and the Gaps

At a glanceQuick answers
Is Step 5 Preview released?
Yes, as an API and product preview. StepFun announced it at 03:15 UTC on September 20, 2026; the model id step-5-preview is documented and priced. Open weights are promised for October 15 and had not appeared on Hugging Face at 13:05 UTC.
What is Step 5 Preview?
StepFun’s flagship model for agentic work: a 600B-parameter sparse mixture of experts with 27B active per token, a 1M-token context window, 64k output tokens, and text, image and video input. StepFun highlights software engineering, long-horizon agent tasks, professional knowledge work and finance.
What does it cost?
On StepFun’s own API, $1.00 per million input tokens on a cache miss, $0.05 on a cache hit, and $2.70 per million output tokens, with reasoning tokens counted as output. Third-party prices did not exist at publication; OpenRouter had no row.
How does it compare?
On StepFun’s table it runs level with Kimi K3 and GLM 5.3 on most rows and behind GPT-6 Astra and Claude Opus 5 on the hardest coding and agent rows. The comparison uses StepFun’s High mode against the rivals’ Max modes, which is the first thing to know before reading any number.
When do the weights land?
StepFun says October 15, 2026. No license is named on the announcement, and no technical report is linked. Until then this is a closed preview with an open-weight promise.
Editorial data illustration on a near-white ground: a curved Pareto frontier line in navy sweeps across a chart of intelligence against cost, an older frontier sits behind it in pale grey, and a rust dot labelled STEP 5 PREVIEW pushes the new curve outward; three huge numbers read 600B with 27B ACTIVE beneath, $2.70 PER 1M OUT, and OCT 15 WEIGHTS
Fig 0A 600B model with 27B awake per token, priced at $2.70 a million out. The frontier moved on cost, not on the top line.

StepFun released Step 5 Preview on September 20, 2026, at 03:15 UTC, and the number that matters is $2.70. That is the price of a million output tokens on StepFun’s own API for a 600B-parameter mixture of experts with 27B active per token, a 1M-token context window, video input, and a benchmark table that puts it beside Kimi K3 and GLM 5.3 on most rows. The announcement is titled “Advancing the Pareto Frontier” and it means that literally: the claim is not the top score, it is the same intelligence for a third of the cost. The weights are promised for October 15. This is the record of what StepFun published, what it measured, and what it has not shipped yet. Times are UTC.

On this page · 9 sectionsOpen
  1. What StepFun shipped
  2. The price is the story
  3. Benchmarks, from StepFun’s own table
  4. The long-horizon runs
  5. What is not here yet
  6. Where CellCog stands
  7. What we are watching for
  8. Update log
  9. Sources
Key points7 · 11 min full read
  1. A calendar page marked September 20 beside an API key card, standing for a same-day release.
    StepFun announced Step 5 Preview at 03:15 UTC on September 20, 2026 (11:15 Beijing), on X and on stepfun.com. It is live in StepFun’s products and API as step-5-preview; the company says the weights follow on October 15.
  2. A large cube with a small lit core, standing for 600B parameters with 27B active.
    The spec, from StepFun’s docs: a sparse mixture of experts with 600B total parameters and 27B active per token (4.5%), a 1M-token context window, 64k output tokens, text, image and video input, three reasoning-effort levels, tool calling, JSON Schema output and prompt caching.
  3. A price tag reading 2.70 next to a much taller tag, standing for the output price gap.
    The price is the story: $1.00 per million input tokens on a cache miss, $0.05 on a cache hit, $2.70 per million output tokens, reasoning included. That is 18% of Kimi K3’s $15.00 output price and 2.35 times StepFun’s own Step 3.7 Flash.
  4. Five bars of different heights with the fourth highlighted, standing for a benchmark comparison.
    StepFun’s own table puts it close to Kimi K3 and GLM 5.3 and behind GPT-6 Astra and Claude Opus 5: DeepSWE v1.1 67.7 (Astra 74.1, Opus 5 74.0), Terminal-Bench v4 33.3 (Astra 57.9), GDPval-AA v2 1571 (Opus 5 1735). It scores 44 on the Artificial Analysis Intelligence Index, per StepFun.
  5. A magnifying glass over a footnote mark, standing for the caveats under a benchmark table.
    The comparison is High against the rivals’ Max modes, six benchmarks are StepFun’s own, and the HLE-with-tools row mixes a text-only subset with full-dataset runs; StepFun says those results are not directly comparable. Read the table with those three notes attached.
  6. A stopwatch showing 24 hours beside a rising line, standing for a day-long autonomous run.
    Three long-horizon runs are the most interesting evidence: 24 hours optimizing an H100 MLA kernel to 508 TFLOPS (Opus 5: 493), 24 hours of automated post-training lifting a Qwen3-30B-A3B base from 53.3% to 60% on AIME24, and 3,000-plus turns of Pokemon Red, a third of the story.
  7. An empty shelf with a date label reading October 15, standing for weights not yet released.
    What is not here yet: no weights, no license named, no technical report, no StepFun repository on Hugging Face, no OpenRouter row. The page calls it a preview and dates the open weights October 15, 25 days out. This record updates the day they land.

§ 01What StepFun shipped

StepFun’s first post, at 03:15:13 UTC, called it “our new flagship model for agentic work, delivering frontier-level performance across software engineering and professional knowledge work, with particular strength in finance.” The developer docs went live with it.

Item Step 5 Preview
Architecture Sparse mixture of experts, 600B total parameters, 27B active per token
Context window 1M tokens input, 64k tokens maximum output
Input Text, images (up to 60 per request) and video (MP4, QuickTime, Matroska)
Reasoning Effort levels low, medium and high, set per request
Features Tool calling, streaming, JSON Mode and JSON Schema, prompt caching
Model id step-5-preview, on an OpenAI-compatible chat completions endpoint and a Messages API
Weights Not released; StepFun says October 15, 2026
License Not named
Technical report None linked
Table 1Step 5 Preview, as documented by StepFun on September 20, 2026

The 27B active figure is the smallest of any model in StepFun’s own comparison table, and it is what the pricing rests on. For scale, Kimi K3 activates 104B of 2.8T; DeepSeek V4.1 Flash activates 8B on input and 16B on output of 552B. Step 5 Preview sits between them in size and, as the next section shows, well below one of them in price.

§ 02The price is the story

Model Input, cache miss Input, cache hit Output
step-5-preview $1.00 $0.05 $2.70
step-3.7-flash $0.20 $0.04 $1.15
Table 2StepFun’s list prices per million tokens, read September 20, 2026

Output tokens include the model’s reasoning, so the $2.70 covers thinking as well as the answer. StepFun’s page adds the framing: “At a comparable level of intelligence, its task cost is substantially lower than similarly capable models,” and gives its Artificial Analysis Intelligence Index score as 44. Against the models StepFun compares itself to, the output price reads like this.

Model Output per 1M tokens Source
GLM-5.3-Flash $0.50 Z.ai’s endpoint on OpenRouter, read September 20
DeepSeek V4.1 Flash $0.60 DeepSeek’s off-peak list price, September 10 notice
Step 3.7 Flash $1.15 StepFun pricing page
Step 5 Preview $2.70 StepFun pricing page
Kimi K3 $15.00 Moonshot’s pricing page and its own OpenRouter endpoint
Table 3Output price per million tokens, each vendor’s own list price
Output price per million tokens, each from the vendor's own pricing pageBar chart of output price per million tokens: GLM-5.3-Flash 0.50, DeepSeek V4.1 Flash 0.60, Step 3.7 Flash 1.15, Step 5 Preview 2.70 highlighted, Kimi K3 15.00GLM-5.3-Flash0.50DeepSeek V4.1 Flash0.60Step 3.7 Flash1.15Step 5 Preview2.70Kimi K315.00Output price per million tokens, each from the vendor's own pricing pageBar chart of output price per million tokens: GLM-5.3-Flash 0.50, DeepSeek V4.1 Flash 0.60, Step 3.7 Flash 1.15, Step 5 Preview 2.70 highlighted, Kimi K3 15.00GLM-5.3-Flash0.50DeepSeek V4.1 Flash0.60Step 3.7 Flash1.15Step 5 Preview2.70Kimi K315.00
Fig 1Output price per million tokens, each from the vendor's own pricing page

The flash-class models are cheaper. The point is the row above them: Step 5 Preview is priced as a flagship, and its flagship comparison set on the benchmark page is Kimi K3 at $15.00 and GLM 5.3, GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5, none of which sells output anywhere near $2.70. Whether the intelligence is comparable is the next section’s question; the cost claim is simply true on the list prices.

§ 03Benchmarks, from StepFun’s own table

StepFun published a 39-row table. The selection below keeps the rows that matter to agent builders, in StepFun’s numbers, and drops Fable 5.1, which StepFun left blank on most of them.

Benchmark Step 5 Preview GLM 5.3 Kimi K3 GPT-6 Astra Claude Opus 5
GPQA Diamond 93.5% 91.7% 93.5% 96.1% 93.2%
HLE 46.5% 42.3% 46.9% 54.7% 54.9%
DeepSWE v1.1 67.7% 66.9% 67.5% 74.1% 74.0%
Terminal-Bench v4 33.3% 41.9% 12.6% 57.9% 52.3%
SWE-Marathon v1.1 72.7% 67.4% 84.4% 77.3% 85.6%
StepCodeBench (StepFun’s own) 49.0% 40.2% 43.9% 61.0% 63.9%
GDPval-AA v2 1571 1634 1548 1580 1735
BrowseComp 88.7% not reported 91.2% 91.5% 90.2%
FrontierFinance 66.4% 64.1% 62.6% 55.0% 69.7%
Agents’ Last Exam (ALE-CLI) 29.5% 28.6% 27.6% 33.3% 28.6%
MMMU-Pro 76.0% text only 81.0% 87.0% 85.0%
Scroll to compare all columns
Table 4Selected rows from StepFun’s benchmark table, Step 5 Preview (High) vs rivals (Max)
DeepSWE v1.1 as reported in StepFun's own table, High vs the rivals' MaxBar chart of DeepSWE v1.1 scores from StepFun's table: GLM 5.3 66.9, Kimi K3 67.5, Step 5 Preview 67.7 highlighted, Claude Opus 5 74.0, GPT-6 Astra 74.1GLM 5.366.9Kimi K367.5Step 5 Preview67.7Claude Opus 574.0GPT-6 Astra74.1DeepSWE v1.1 as reported in StepFun's own table, High vs the rivals' MaxBar chart of DeepSWE v1.1 scores from StepFun's table: GLM 5.3 66.9, Kimi K3 67.5, Step 5 Preview 67.7 highlighted, Claude Opus 5 74.0, GPT-6 Astra 74.1GLM 5.366.9Kimi K367.5Step 5 Preview67.7Claude Opus 574.0GPT-6 Astra74.1
Fig 2DeepSWE v1.1 as reported in StepFun's own table, High vs the rivals' Max

The pattern is consistent: level with the two Chinese open-weight flagships, six to seven points behind Astra and Opus 5 on the hardest coding rows, ahead of everyone on FrontierFinance except Opus 5, and weak on multimodal document work (GDP.pdf 14.8% against Astra’s 31.0%). Three caveats, all from StepFun’s own footnotes. The comparison pits Step 5 Preview’s High effort against the rivals’ Max modes. Six rows carry a dagger for benchmarks StepFun developed internally, StepCodeBench and the three FinStepBench suites among them. And the HLE-with-tools row runs Step 5 and GLM 5.3 on a text-only subset against the full dataset for the others; StepFun writes that “Results across these evaluation settings are not directly comparable.” We agree, and we left that row out.

StepCodeBench deserves its own line because StepFun built it: 553 independent repositories, 9 task categories, 20 application domains, 33 programming languages, scored avg@4. StepFun reports 49.0% and says the model does best on bug repair, feature modification and refactoring, with “a meaningful gap to the frontier” on the longest tasks. That is a more useful admission than most launch pages make.

§ 04The long-horizon runs

The most interesting evidence on the page is not a benchmark. StepFun ran three open-ended tasks with a clock and a scoreboard.

Task Budget Result Comparison StepFun gives
Optimize an MLA GPU kernel from scratch on an H100 (head dim 512, batch 1, 64 heads, 8,192 tokens) 24 hours, best of 4 attempts 508 TFLOPS forward and backward after roughly 22 hours Claude Opus 5: 493 TFLOPS
Improve a Qwen3-30B-A3B base on AIME24 through automated post-training with an API annotator 24 hours 60% on the official test, up from 53.3% Matched Opus 5 using fewer annotator tokens
Play Pokemon Red with no game-specific optimization Open 3,000-plus turns, 6 million tokens, three Gym Badges by turn 3,082 About a third of the main story
Table 5StepFun’s three long-horizon experiments, as reported

A 3% kernel throughput edge over Opus 5 in a vendor-run experiment is a data point, not a verdict; the AIME24 run matched rather than beat. What the three runs show together is the property StepFun is actually selling: a model that keeps going for a day, checks its own results and discards regressions, at a price where a day of that is affordable. The same page reports one agent action coordinating 950 web fetches for a 1,000-location climate study, and “approximately 70% of participants” among internal and external experts judging the model able to solve moderately complex coding tasks on its own.

§ 05What is not here yet

Item Status
API access Live, model id step-5-preview, priced
Open weights Promised for October 15; none on Hugging Face under stepfun-ai
License Not named anywhere on the announcement or docs
Technical report None linked
Third-party hosting No OpenRouter row; only Step 3.5 Flash and Step 3.7 Flash listed
Independent evaluation StepFun cites its Artificial Analysis Intelligence Index score as 44; the AA page is the place to confirm it
Launch event or livestream None; a web post and a six-post X thread
Table 6Step 5 Preview: claimed vs shipped, September 20, 2026, 13:05 UTC

§ 06Where CellCog stands

Our conflict, declared: we build CellCog, a platform where a business hires AI employees that keep memory and work as a team, and we choose the models underneath them. Today that is Claude Fable 5.1 at Core and Max and Gemini 3.8 Flash at Flash. We do not route to StepFun. We read Step 5 Preview the way an operator reads it: a candidate engine whose price makes day-long autonomous runs cheap enough to try, and whose open-weight release on October 15 will decide whether it can be run on our own terms. If the license allows it, it gets tested on our infrastructure that day.

§ 07What we are watching for

  • October 15. A stepfun-ai repository on Hugging Face with a named license. The announcement’s date is the claim; the repository is the fact.
  • A third-party row. OpenRouter or another host listing step-5-preview with its own price, which is where a vendor’s list price meets the market.
  • The technical report. None is linked; parameter count, training tokens and the MoE routing are described in one paragraph.
  • Independent numbers. Artificial Analysis publishing its own Step 5 Preview index page, and the SWE-agent harness runs behind the DeepSWE row being reproduced by anyone else.
  • A mode-matched comparison. Any table that runs Step 5 Preview at High against rivals at their equivalent effort rather than Max.

§ 08Update log

This is a living page; when the story moves, the update lands here.

As of September 20, 2026, 13:10 UTC: page opened. StepFun’s X thread read in a browser and clocked from post ids (03:15:13 UTC); the announcement page read in a browser at 13:03 UTC; the model, quickstart and pricing docs read from platform.stepfun.ai at 13:02 UTC; Hugging Face and OpenRouter read at 13:01 and 13:05 UTC, no Step 5 repository or row.

§ 09Sources

Frequently asked5 questions

Q1Is Step 5 Preview open source?

Not yet. StepFun’s announcement says the model will be released with open weights on October 15, 2026, and does not name a license. As of 13:05 UTC on September 20 the stepfun-ai organization on Hugging Face has no Step 5 repository; user-owned repositories carrying the name are not StepFun’s.

Q2Can I run it through OpenRouter or another host?

Not at publication. OpenRouter listed two StepFun models, Step 3.7 Flash and Step 3.5 Flash, and no step-5-preview id when we read the catalog at 13:01 UTC. StepFun’s own API at api.stepfun.ai serves it under the OpenAI-compatible chat completions endpoint and, per the docs, a Messages API.

Q3How does it compare to Kimi K3?

Kimi K3 is a 2.8T-parameter model with 104B active and open weights since July 27; Step 5 Preview is 600B with 27B active and closed until October 15. On StepFun’s table the two trade rows: K3 ahead on SWE-Marathon (84.4 vs 72.7) and BrowseComp (91.2 vs 88.7), Step 5 ahead on Terminal-Bench v4 (33.3 vs 12.6) and StepCodeBench (49.0 vs 43.9). On price, Step 5’s $2.70 output is 18% of K3’s $15.00 on Moonshot’s own API.

Q4What are the caveats in StepFun's benchmark table?

Three, all stated on the page. Step 5 Preview ran at its High effort while GPT-6 Astra, Claude Opus 5, Claude Fable 5.1, Kimi K3 and GLM 5.3 ran at Max. Six rows marked with a dagger are StepFun’s own benchmarks, including StepCodeBench and the three FinStepBench suites. The HLE-with-tools row compares a text-only subset for Step 5 and GLM 5.3 against the full dataset for the others, which StepFun itself calls not directly comparable.

Q5Should I move my agents to Step 5 Preview?

Test it if output cost is your constraint and your tasks look like StepFun’s strong rows: bug repair, refactoring, research with many tool calls, finance. Hold if you need open weights, a stable identifier, or the top line on the hardest coding benchmarks. If your agents run on CellCog, the engine is a setting: we run Claude Fable 5.1 at Core and Max and Gemini 3.8 Flash at Flash today, we do not route to StepFun, and we would test Step 5 the day its weights land.

Published 20 September 2026 All Choosing a platform →