StepFun (阶跃星辰)
StepFun (阶跃星辰) is a Chinese AI company developing multimodal foundation models and speech models. The corporate entity is Shanghai Jieyue Xingchen Intelligent Technology Co., Ltd. (上海阶跃星辰智能科技股份有限公司), founded in Shanghai on April 6, 2023 by Jiang Daxin, Zhu Yibo, and Jiao Binxing, all with Microsoft backgrounds (per English Wikipedia).
Models and developer platform
StepFun’s models carry the Step name: the trillion-parameter Step-2 LLM and multimodal models arrived in July 2024; in February 2025 the company open-sourced Step-Video-T2V and the Step-Audio speech models with Geely; Step 3 followed in July 2025; and Step 3.5 Flash, released in February 2026, is a mixture-of-experts model under the Apache 2.0 license (timeline per English Wikipedia). On the speech side the lineup has grown into the Step-Audio 3 family, covering real-time interaction, recognition, speech generation, and music generation.
Developers can use the open platform: one API key covers language, speech, and image models, and the Chat Completions endpoint is compatible with the OpenAI SDK; the platform also offers AI Studio for browser-based testing and programs for startups.
Discussion in the show
In Next Token Weekly #003’s chapter “用 GPT-6 Astra 造数据、训练小模型” (training small models with GPT-6 Astra), Guizang reports that “StepFun released three speech models — the real-time StepAudio 3 Realtime, StepAudio 3 ASR Max, and full audio content generation”. See the Step-Audio entry and the episode 003 chapter in the Chinese transcript.
Frequently asked questions
What company is StepFun?
StepFun (阶跃星辰) is an AI company founded in Shanghai in April 2023 that develops the Step series of multimodal foundation models and the Step-Audio speech model family. Its founders are Jiang Daxin and two co-founders, all formerly at Microsoft.
Where are StepFun’s website and developer platform?
The official website is stepfun.com, and the developer platform is platform.stepfun.com, with API documentation, key management, and model playgrounds.
What is Step-Audio, and how does it relate to StepFun?
Step-Audio is StepFun’s speech model family, covering real-time voice interaction, recognition, and speech generation. It was open-sourced with Geely in February 2025 and has since iterated into the Step-Audio 3 family. See the Step-Audio entry.
Who founded StepFun?
Per English Wikipedia, StepFun was co-founded in 2023 by Jiang Daxin, Zhu Yibo, and Jiao Binxing, all of whom previously worked at Microsoft.