StepFun launches a 600B agent model at $1 per million input tokens
Jiang Daxin's flagship pairs a 1M-token window with aggressive pricing, while its own tests still trail top rivals on several coding benchmarks.
By Ryan Merket · Published
Primary source: StepFun
Why it matters
StepFun is pairing open-weight models with a proprietary flagship, using low agent costs as its wedge while its own benchmarks still show a capability gap.

StepFun founder and CEO Jiang Daxin introduced Step 5 Preview, betting that a sparse 600-billion-parameter model can win agent workloads with low operating costs despite trailing top rivals on several company-reported benchmarks.
StepFun says Step 5 Preview activates 27 billion parameters for each token, accepts images and text, and supports a 1-million-token context window. The proprietary model is aimed at software engineering, professional knowledge work and finance, with tools for executing longer jobs that require repeated testing and revision.
StepFun was founded in 2023 and is headquartered in Shanghai, according to StepFun's public profile. Its developer platform lists an address in the city's Xuhui district. StepFun sells language, voice, image and multimodal models through APIs and also offers consumer AI products across web, desktop and mobile devices.
The release brings Jiang back to the field that shaped much of his career. Before founding StepFun, he spent 16 years at Microsoft, where his work covered Bing search, Cortana, Azure Cognitive Services and natural-language systems. He previously earned a computer science Ph.D. from the University at Buffalo.
ChatGPT supplied the trigger. After OpenAI released it in November 2022, Jiang recalled thinking, "I can do it myself, maybe even better," according to a 2025 profile of StepFun. Step 5 Preview is his clearest attempt yet to turn that conviction into a model positioned alongside the largest commercial systems.
The price is the product
StepFun's launch materials describe a shift in the "Pareto frontier" between intelligence and cost. The independent figures available so far support the cost portion of that pitch.
Artificial Analysis gives Step 5 Preview a score of 44 on its Intelligence Index and lists API pricing of $1 per million input tokens and $2.70 per million output tokens. Artificial Analysis calculated a weighted cost of $0.71 per Intelligence Index task. Its page also lists a 95% cache discount, which could matter for agents repeatedly working from the same codebase or document set.
That pricing sits below Artificial Analysis' medians of $1.88 for input and $10 for output among comparable models. Step 5 Preview generated 160 million output tokens during the evaluation, compared with a 92-million median for comparable reasoning models. A model that reasons at length can consume some of the savings created by a low per-token price, although Artificial Analysis' task-level calculation still places Step 5 Preview among the better-priced models at its reported capability level. Artificial Analysis reports an output speed of approximately 99.8 tokens per second.
Sparse activation is central to Jiang's economics. Step 5 Preview uses 4.5% of its total parameters for each token. That allows StepFun to advertise the capacity of a 600-billion-parameter model without paying the full inference cost on every generation, although actual deployment expenses also depend on memory movement, serving infrastructure, prompt caching and how long the model reasons.
StepFun's benchmark table is a mixed result
StepFun reports that Step 5 Preview scored 67.7 on DeepSWE v1.1, 49 on StepCodeBench and 80.5 on ProgramBench. Those results put it ahead of Kimi K3 and GLM-5.3 in StepFun's table, while GPT-6 Astra and Claude Opus 5 remained ahead on all three tests.
StepCodeBench is StepFun's own evaluation. The company says it covers 553 independent repositories, nine task categories, 20 application domains and 33 programming languages. The breadth makes the reported score more informative than a narrow single-language test, though buyers still have to rely on StepFun's construction and administration of the benchmark.
The gap widened on Terminal-Bench v4. StepFun reported a score of 33.3 for Step 5 Preview, compared with 41.9 for GLM-5.3, 52.3 for Claude Opus 5 and 57.9 for GPT-6 Astra. On Agents' Last Exam, Step 5 Preview's reported 29.5 exceeded Claude Opus 5 and the two other Chinese models in the table, while remaining below GPT-6 Astra's 33.3.
Finance produced the more favorable comparison that StepFun emphasizes. Step 5 Preview scored a company-reported 66.4 on FrontierFinance, ahead of GPT-6 Astra's 55 and below Claude Opus 5's 69.7. On DRACO, another finance evaluation, StepFun reported 83.3, compared with 76.8 for GPT-6 Astra and 87.6 for Claude Opus 5.
StepFun also published two 24-hour agent experiments. In one, the company says Step 5 Preview optimized an inference kernel to 508 TFLOPS on an Nvidia H100 after roughly 22 hours, compared with 493 TFLOPS for Claude Opus 5. In another, the model managed an automated post-training loop that raised a Qwen3-30B-A3B base model's AIME24 accuracy from 53.3% to 60%, matching StepFun's reported result for Claude Opus 5 while using fewer annotator tokens.
These are StepFun-published comparisons rather than a common set of independently reproduced results. StepFun identifies the SWE-agent harness and sampling settings used for DeepSWE, but the release does not provide equivalent methodological detail for every score. StepFun also says roughly 70% of internal and external experts judged the model capable of autonomously completing coding tasks of moderately high complexity. The participant count, task set and grading procedure are not included on the release page, limiting the value of that percentage for buyers comparing production systems.
Artificial Analysis offers the cleaner third-party check. Its score of 44 places Step 5 Preview 27th among 653 models on the current index, according to its model page. That is a strong debut at StepFun's price point and a less sweeping result than the phrase "frontier-level performance" on the launch page suggests.
Jiang is running an open and closed model strategy
Step 5 Preview also clarifies how Jiang plans to divide StepFun's model lineup. On February 12th, StepFun released Step 3.5 Flash, an open-weight mixture-of-experts model with 196 billion total parameters, 11 billion active parameters and a 256,000-token context window. StepFun promoted that model for local deployment, coding agents and high-throughput inference.
Step 5 Preview is larger, proprietary and priced as an API service. The pairing gives StepFun an open model that developers can deploy and customize, alongside a closed flagship that StepFun can meter and monetize. Artificial Analysis currently lists Step 5 Preview as its new flagship alongside Step 3.7 Flash and Step 3.5 Flash, while StepFun's public developer platform presents Step 5 Preview as its new flagship and lists it alongside Step 3.7 Flash and Step 3.5 Flash.
The platform says registered developers can call all public models through a single API key and pay according to usage. It also publishes an OpenAI-compatible endpoint, reducing the integration work for teams that already use that client format. Public rate limits, geographic restrictions and enterprise access terms for Step 5 Preview have not been established in the available materials.
Jiang has substantial capital behind the push. StepFun raised more than $718M in a Series B+ round announced on January 26th, according to TechNode. New participants included Shanghai SDIC Leading Fund, China Life Private Equity Investment, Pudong Venture Capital, Xuhui Capital, Wuxi Liangxi Fund, Xiamen ITG Group and Huaqin Technology. Tencent, Qiming Venture Partners and 5Y Capital invested again.
That followed a Series B worth several hundred million dollars in December 2024, according to the South China Morning Post. Fortera Capital, the private-equity arm of Shanghai State-owned Capital Investment, led that round, with Tencent and Qiming Venture Partners also participating. The exact amount and valuation were not disclosed.
The January 2026 round was earmarked for foundation-model development and StepFun's strategy of putting AI into phones, cars and other devices. Step 5 Preview is a larger general-purpose model designed to power agents in the cloud, alongside other models in StepFun's catalog for broader product and device use.
For Jiang, the immediate test is usage rather than another StepFun-generated chart. Step 5 Preview has a credible price advantage, a long context window and competitive results in finance. StepFun's own numbers also show clear deficits on several coding and terminal tasks. Converting inexpensive tokens into reliable completed work will determine whether Jiang has moved the frontier or simply lowered the rate card.