StepFun ships five StepAudio 3 models across voice, sound and music

The API release tops two independent voice benchmarks, while its 8.83-second response time leaves StepFun with a latency problem.

By · Published

Primary source: StepFun on X

Why it matters

StepFun is selling a single API stack for voice agents and media generation. Its quality scores are strong, but latency and preview pricing will decide whether developers stay.

A person wearing sleek headphones listens intently in a modern workspace, surrounded by subtle, flowing visual patterns of sound.

StepFun (@StepFun_ai), the Shanghai AI lab co-founded and led by Daxin Jiang, released five StepAudio 3 models on September 15th, giving developers separate APIs for real-time voice interaction, transcription, speech synthesis, general audio generation and music.

The release packages most of the audio stack required for a voice agent or media-production workflow under one model family. StepAudio 3 Realtime handles full-duplex conversation and tool calls; ASR Max transcribes speech; TTS generates voices; Gen combines dialogue, sound effects, ambience and background music; and Music turns prompts, lyrics, vocals or reference audio into songs and instrumentals.

Jiang founded StepFun in 2023 after a long career at Microsoft, where he worked on speech and Azure AI and served as a global vice president. That background makes StepAudio 3 a direct extension of the founder's earlier work rather than a side product attached to a text-model lab. StepFun has pursued multimodal models from its founding, spanning text, images, video and audio.

The benchmark lead comes with a speed penalty

StepFun's strongest evidence comes from independent testing. Artificial Analysis ranks StepAudio 3 Realtime first for speech reasoning, with 99.7% on its 1,000-question Big Bench Audio test. The model also leads conversational dynamics at 98.9%, narrowly ahead of Alibaba Cloud's Qwen Audio 3.0 Realtime Plus at 98.4%.

Conversational dynamics measures whether a model handles pauses, interruptions, turn-taking and short acknowledgments without talking over the user or ending a response prematurely. Those behaviors tend to determine whether a voice agent feels usable after the demo ends.

The same evaluation exposes StepFun's immediate weakness. StepAudio 3 Realtime took an average 8.83 seconds to produce its first audio on the benchmark. The fastest model listed, Deepslate Opal, started in 0.44 seconds, while several models from Google, OpenAI, Alibaba Cloud and xAI responded in roughly one to three seconds.

StepAudio 3 Realtime therefore leads on measured reasoning and conversation handling while trailing badly on initial response speed. StepFun's technical report, submitted on September 12th, describes a "Think-While-Speaking" system that runs private reasoning alongside spoken output. The design targets the tension between deliberation and latency, although the independent time-to-first-audio result shows that the tension remains unresolved in the tested deployment.

The model also lacks enough published results to qualify for Artificial Analysis' overall speech-to-speech index, which requires scores for reasoning, agentic performance, human preference and task completion. StepFun has two category leads, rather than a clean sweep of the broader evaluation.

Transcription accuracy is tied for first

Artificial Analysis' speech-to-text leaderboard gives StepAudio 3 ASR a 1.7% word error rate across about eight hours of audio covering varied accents, specialist language and difficult recording conditions. That ties Alibaba Cloud's Fun-Realtime-ASR-preview for the lowest error rate among 59 models tested.

StepFun's ASR documentation says the model uses context and domain knowledge to recognize names, homophones and technical vocabulary across fields including medicine, finance, law and software development. StepFun prices ASR Max at $0.24 per audio hour, nearly 11 times the $0.022 hourly price of StepAudio 2.5 ASR.

The pricing difference makes ASR Max a premium accuracy product inside StepFun's own catalog. Customers that need cheap batch transcription can stay on the earlier model, while workloads involving specialist terms or difficult audio carry the higher rate.

StepAudio 3 TTS moves in the opposite direction. The pricing page lists it at $0.36 per 10,000 characters, nearly 58% below the $0.85 price of StepAudio 2.5 TTS. Voice cloning costs another $1.50 per voice.

StepAudio 3 Realtime, Gen and Music are available as free, limited-time preview APIs. StepFun says the preview identifiers will be retired when paid versions arrive, leaving the eventual production pricing for three of the five models unsettled.

One family, two different markets

The five-model release puts StepFun in two crowded businesses at once. Realtime, ASR and TTS target developers building customer-service agents, assistants and voice interfaces. Gen and Music compete for media-production work that currently requires separate voice, sound-effects and music tools.

StepAudio 3 Gen uses a shared discrete audio representation for speech, vocals, sound effects and music, according to a technical report submitted on September 11th. The model predicts compressed audio tokens autoregressively, allowing one prompt to specify voices, timing, ambience and music in a single clip.

That consolidation is StepFun's product bet. A developer can assemble an audio application through one provider instead of connecting separate transcription, reasoning, synthesis and generation services. The benchmark results give StepFun a credible opening on quality. Production adoption will depend on whether StepFun can cut real-time latency and convert the free preview models into pricing that remains competitive once the launch window closes.

Reader comments

Conversation for this story loads after sign-in.