YFarmX logoYFarmX

StepFun

Step 5 Preview

StepFun's flagship for long agent jobs, with open weights due on 15 October

Released 20 September 202612 min readLarge Language Models

Editorial collage on off-white newsprint with torn blue corners: a staircase of five white paper cubes rising left to right, the top step cyan, beneath a large STEP 5 PREVIEW headline and the line STEPFUN FLAGSHIP · API OPEN · WEIGHTS DUE 15 OCT, with the StepFun logo, a torn spec sheet reading 600B TOTAL, 27B ACTIVE and 1M CONTEXT, a price tag reading $1.00 IN · $2.70 OUT and a desk calendar page circled at 15 OCT.

Key facts

20 Sep 2026API, AI Studio, Step Plan
Released
600B27B active per token
Size
1M / 64Ktokens in / out
Context
$1.00 / $2.70per M tokens in / out
Price
44v4.3.2, Artificial Analysis
Intelligence Index
15 Oct 2026StepFun's stated date
Open weights

Step 5 Preview is an AI language model from StepFun, a Shanghai company, built to carry long jobs through on its own: writing and fixing code, gathering research, building financial analysis. It went on sale on 20 September 2026, reads text, images and video, holds a million tokens of context, and costs $1 for every million tokens you send it and $2.70 for every million it writes. Artificial Analysis scores it 44 on its Intelligence Index, level with Kimi K3 and Grok 4.6, and StepFun says the weights will be published for anyone to download on 15 October 2026.

What Step 5 Preview is

Step 5 Preview is a large language model from StepFun, a Shanghai artificial intelligence company, built to work through long jobs on its own: writing and fixing code, gathering and checking research, and building financial analysis. StepFun announced it on 20 September 2026 in a post on X, calling it “our new flagship model for agentic work”, and its launch page says it went on sale the same day “through our products and API”. It reads text, images and video, writes text, and holds a context window of one million tokens, according to StepFun’s model documentation.

The model is a mixture of experts. It stores 600 billion parameters, the numbers it learned in training, and runs 27 billion of them for each token it reads or writes, so each request costs far less to serve than the full size suggests. StepFun gives both figures on its launch page, and Artificial Analysis, the independent benchmarking company, repeats them in its own write-up of 22 September 2026.

StepFun works from Xuhui District in Shanghai, the address on its own site. Its previous reasoning model, Step 3.7 Flash, is a 198B model with open weights on Hugging Face under Apache 2.0. Artificial Analysis dates Step 3.7 Flash to 29 May 2026 and scores it 19 on the same Intelligence Index where Step 5 Preview scores 44.

What it costs

StepFun charges $1.00 for every million tokens sent to Step 5 Preview and $2.70 for every million it writes back, on its pricing page read on 27 September 2026. Repeated input that hits the cache costs $0.05 per million, a 95% discount.

Per million tokens step-5-preview step-3.7-flash
Input, cache miss $1.00 $0.20
Input, cache hit $0.05 $0.04
Output $2.70 $1.15

The cache-miss price already covers writing new content into the cache. Output tokens count the model’s reasoning as well as its final answer, which weighs on this model in particular: to run the Intelligence Index, Step 5 Preview generated 160 million output tokens against a median of 88 million for comparable models, which Artificial Analysis calls “very verbose”. Its measured cost comes to $0.72 per index task.

Pay-as-you-go accounts are rate limited by how much cash they have topped up in total. Only cash top-ups count towards a tier, and individual accounts complete identity verification first.

Tier, by total top-up Concurrent requests Requests a minute Tokens a minute
V0, under $15 5 100 500,000
V1, $15 to $69 20 400 2,000,000
V2, $70 to $299 30 600 3,000,000
V3, $300 to $1,499 40 800 4,000,000
V4, $1,500 and above 130 2,600 13,000,000

You can also pay monthly through Step Plan

Step Plan, StepFun’s subscription for coding tools and agent apps, includes Step 5 Preview from $6.99 a month, according to the Step Plan overview. Each tier issues a monthly allowance of Credits, which StepFun converts at about 7 million Credits to $1 of model use, and unused Credits clear at the end of the month.

Plan Monthly Credits Monthly Yearly
Flash Mini 400M $6.99 $69.99
Flash Plus 1,600M $9.99 $95.99
Flash Pro 8,000M $29 $289
Flash Max 40,000M $99 $989

Quarterly billing costs $18.99, $26.99, $79 and $269 for the four tiers. At StepFun’s conversion rate, Flash Mini’s 400 million Credits buy roughly $57 of usage for $6.99. Step Plan is exempt from the top-up rate limits, and booster packs at $6.99 for 400M Credits and $9.99 for 1,600M top up a month that runs dry. StepFun publishes setup guides for Claude Code, Cursor, Cline, OpenClaw and eleven more tools, and its Claude Code guide switches on the full one-million-token context.

Step 5 Preview ties Kimi K3 on the independent index

Artificial Analysis scores Step 5 Preview 44 on version 4.3.2 of its Intelligence Index, level with Kimi K3 (max) and Grok 4.6 (high), on its model pages read on 27 September 2026. The index combines ten evaluations run by Artificial Analysis itself: AA-Briefcase, GDPval-AA, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR.

Artificial Analysis, 27 September 2026 Intelligence Index v4.3.2 Cost per index task Input / output, per M tokens
Step 5 Preview 44 $0.72 $1.00 / $2.70
Kimi K3 (max) 44 $2.00 $3.00 / $15.00
Grok 4.6 (high) 44 $1.86 $2.00 / $6.00
GLM-5.3 (max) 45 $2.01 $1.40 / $4.40
MiMo-V2.6-Pro 46 $0.13 $0.435 / $0.87

In its post of 22 September 2026, Artificial Analysis says Step 5 Preview is “matching Kimi K3 (max) at ~2.8x lower cost per task, but trails peers on agentic evaluations”. Against Kimi K3, Grok 4.6 and GLM-5.3, it does the index’s work at a little over a third of their cost per task, and its evaluation scores split as follows.

Evaluation Step 5 Preview Best of the other three Whose
Terminal-Bench 4.0 33.3% 41.9% GLM-5.3 (max)
AutomationBench-AA 51.0% 66.7% Grok 4.6 (high)
AA-Briefcase v1.1 (Elo) 1,432 1,539 Grok 4.6 (high)
AA-LCR v1.1, long context 88.3% 88.7% Kimi K3 (max)
Humanity’s Last Exam 46.5% 46.9% Kimi K3 (max)

Step 5 Preview scores lowest of the four on AutomationBench and AA-Briefcase, the two tests of multi-step business work, where Kimi K3 scores 58.3% and 1,505 and GLM-5.3 62.2% and 1,517. It handles long documents about as well as Kimi K3, and on Terminal-Bench 4.0 it clears Kimi K3’s 12.6% and Grok 4.6’s 21.2%.

On StepFun’s own API, Artificial Analysis measures 66.4 output tokens a second, slower than the average of 78 across the models it compares it with, and 2.96 seconds to the first token.

Artificial Analysis Intelligence Index v4.3.2 bar chart of 28 models, with Step 5 Preview highlighted in cyan at 44, between GLM-5.3 (max) at 45 and Gemini 3.8 Flash (high) at 41; Claude Opus 5.5 leads at 58
Step 5 Preview at 44 on the Intelligence Index, in the default 28-model view, one point behind GLM-5.3 (max). Screenshot of Artificial Analysis's Step 5 Preview page, 27 September 2026.

Xiaomi’s MiMo V2.6 Pro scores higher for less

Xiaomi’s MiMo V2.6 Pro, released on 21 September 2026, scores 46 on the Intelligence Index at $0.13 a task on Artificial Analysis, against Step 5 Preview’s 44 at $0.72. Among the four models scoring 44 or 45, Step 5 Preview, Kimi K3, Grok 4.6 and GLM-5.3, Step 5 Preview is the cheapest per task by a wide margin. StepFun titled its launch “Advancing the Pareto Frontier”, the line of best score for the money, and on Artificial Analysis’s per-task costs Step 5 Preview’s $0.72 is about 64% below GLM-5.3’s $2.01 and Kimi K3’s $2.00. MiMo V2.6 Pro arrived the next day with open MIT-licensed weights at 1.02 trillion parameters, and Artificial Analysis’s cost chart now draws its frontier line through MiMo, above and to the left of Step 5 Preview.

The comparison splits by what a buyer needs. Against Kimi K3, which Artificial Analysis lists at 2.8 trillion parameters and $2.00 a task, Step 5 Preview reaches the same index score at a little over a third of the cost. Against GLM-5.3 at $2.01 a task, it gives up one index point and 8.6 points of Terminal-Bench 4.0 for a similar saving. Grok 4.6 scores 44 at $1.86 a task, leads Step 5 Preview on AutomationBench and AA-Briefcase, and trails it on Terminal-Bench 4.0.

StepFun’s own tests put it level with Kimi K3 and behind Claude

On StepFun’s own runs, published on its launch page on 20 September 2026, Step 5 Preview beats GLM-5.3 (max) on 25 of 38 benchmarks, splits Kimi K3 (max) almost evenly at 18 wins, 19 losses and 3 ties across 40, and trails Claude Opus 5 (max) on 33 of 38. A selection of StepFun’s figures:

Benchmark, StepFun’s figures Step 5 Preview (high) Kimi K3 (max) Claude Opus 5 (max)
GPQA Diamond 93.5% 93.5% 93.2%
DeepSWE v1.1 67.7% 67.5% 74.0%
ProgramBench, pass rate 80.5% 77.8% 82.3%
SWE-Atlas-QnA 63.6% 59.5% 66.0%
Toolathlon-Verified 74.1% 76.5% 80.6%
BrowseComp 88.7% 91.2% 90.2%
FrontierFinance 66.4% 62.6% 69.7%
StepCodeBench, StepFun’s own test 49.0% 43.9% 63.9%

StepCodeBench is StepFun’s in-house coding test, built from 553 repositories across 9 task categories, 20 application domains and 33 programming languages, and Step 5 Preview scores 49.0% averaged over four attempts. StepFun describes it as “approaching Claude Opus 5 on low- and medium-difficulty tasks”, and says that “on the most challenging long-running tasks, a meaningful gap to the frontier remains”. DeepSWE v1.1 ran on the SWE-agent harness at a temperature of 1.0 and a top_p of 0.95.

Finance is the domain StepFun singles out. FrontierFinance is an outside benchmark of 220 expert-written questions scored against 11,543 criteria, and StepFun’s three in-house FinStepBench tests cover live financial search, company valuation and full research reports.

Step 5 Preview spent 24 hours tuning a GPU kernel

StepFun gave Step 5 Preview 24 hours to speed up an MLA attention kernel, a piece of low-level GPU code, on an NVIDIA H100, and reports that its best run reached 508 TFLOPS across the forward and backward passes, against 493 for Claude Opus 5. The kernel used a head dimension of 512 at a production shape of batch size 1, 64 heads and 8,192 tokens. Each model had four independent attempts, free to stop early, and StepFun reports the best of the four. In the same test Kimi K3 reached 307 TFLOPS and GLM-5.3 286, by StepFun’s count.

StepFun line chart titled MLA-512 throughput vs solving time, plotting achieved TFLOPS against 24 hours of solving time: Step 5 Preview (High) climbs to 508, Claude Opus 5 (Max) to 493, Kimi K3 (Max) to 307 and GLM-5.3 (Max) to 286, with dots marking each new record and crosses marking attempts
Each line is a model's best kernel so far, dots mark new records and crosses mark attempts that ran slower. Chart from StepFun's Step 5 Preview launch page, 20 September 2026.

StepFun ran two more long jobs. In a second 24-hour test, Step 5 Preview ran automated post-training on a Qwen3-30B-A3B base model and lifted its AIME24 score from 53.3% to 60%, which StepFun says matched Claude Opus 5 using fewer annotator tokens. And in Pokémon Red it had played more than 3,000 turns and 6 million tokens by the time the launch page went up: by turn 3,082 it held three Gym Badges and had beaten Lt. Surge.

What you can send it

Step 5 Preview takes up to 1 million tokens in and writes up to 64K tokens out per request, according to StepFun’s model documentation.

Item Supported
Model ID step-5-preview
Input Text, images and video
Output Text, up to 64K tokens
Images JPG, PNG, WebP and static GIF, up to 60 a request, by URL or Base64, at low or high detail
Video MP4, QuickTime and Matroska, by URL, Base64 or a Files API reference; MP4 files under 128 MB, under 5 minutes recommended
Reasoning effort low, medium or high
Also supported Streaming, tool calling, JSON Mode and JSON Schema, prompt caching

Reasoning effort is set with reasoning_effort on the OpenAI-style Chat Completions API and with output_config.effort on the Anthropic-style Messages API. Search, code execution and other tools come from the application calling the model, which reaches your files and outside services only through them.

How to try it

Step 5 Preview runs in three StepFun products. The API at api.stepfun.ai takes the model ID step-5-preview. Step Plan carries it into coding tools such as Claude Code and Cursor. And AI Studio, StepFun’s browser app, runs it with the thinking level and sampling settings in a side panel; on 27 September 2026 it was the model selected for Generate when the app opened.

StepFun AI Studio in a web browser: a prompt box headed Where would you like to start today?, and a Generate parameters panel on the right with Model set to Step 5 Preview, Thinking set to Medium, temperature 1.00 and top_p 0.95
AI Studio with Step 5 Preview selected for Generate, thinking at Medium and sampling at StepFun's defaults. Screenshot of StepFun AI Studio, 27 September 2026.

Artificial Analysis lists one API provider for the model, StepFun itself, and its speed and latency figures come from that first-party endpoint.

Open weights are due on 15 October

StepFun’s launch page says Step 5 Preview will be released with open weights on 15 October 2026, and the StepFun homepage runs a countdown to the release, which on 27 September 2026 read 17 days. Until that date the model runs only on StepFun’s own services, and Artificial Analysis classes it as proprietary.

Step 3.7 Flash, StepFun’s previous reasoning model, went out under Apache 2.0, which allows commercial use. The licence StepFun attaches to Step 5 Preview’s weights decides whether other providers can serve the model commercially.

What to watch

15 October and the licence. StepFun set the weights date itself, and the licence decides whether other hosts pick the model up and compete on price.

The agent scores. Artificial Analysis says Step 5 Preview trails its peers on agent evaluations, and AutomationBench and AA-Briefcase are where the gap is widest. A model named Preview leaves room for a later build that closes it.

The token bill. Step 5 Preview wrote 160 million output tokens to finish one index run, nearly twice the median, so a real bill depends as much on how long it thinks as on its price per token. Measure a real workload before committing to it.

For the models StepFun measures itself against, see Kimi K3 and GLM-5.3; for the cheaper open model that arrived a day later, MiMo V2.6.

Questions people ask

How much does Step 5 Preview cost?
StepFun charges $1.00 per million input tokens, $0.05 per million on a cache hit and $2.70 per million output tokens, on its pricing page read on 27 September 2026. Output tokens include the model's reasoning as well as its answer. Step Plan, StepFun's monthly subscription, includes the model from $6.99 a month.
How big is Step 5 Preview?
Step 5 Preview is a mixture-of-experts model with 600 billion parameters in total, of which 27 billion run for each token. StepFun gives both figures on its launch page of 20 September 2026, and Artificial Analysis repeats them in its own write-up of 22 September.
What does Step 5 Preview score on the Artificial Analysis Intelligence Index?
It scores 44 on version 4.3.2 of the index, level with Kimi K3 (max) and Grok 4.6 (high), as of 27 September 2026. Artificial Analysis puts its cost at $0.72 per index task, against $2.00 for Kimi K3 (max) and $1.86 for Grok 4.6 (high).
When will Step 5 Preview's weights be released?
StepFun's launch page says the model will be released with open weights on 15 October 2026, and its homepage runs a countdown to the date. Until then it runs on StepFun's own API, its AI Studio and its Step Plan subscription, and Artificial Analysis lists StepFun as its one API provider.