Fish Audio adds speaker and emotion labels to its speech transcripts

The feature is available through Fish Audio's transcribe-1-pro API at $0.36 per audio hour; the company says its speech-recognition models support 83 languages.

By · Published

Primary source: X

Why it matters

Fish Audio is extending a voice-generation business into speech recognition, where speaker and vocal cues can make transcripts more useful to downstream software. The feature is available at a stated API rate of $0.36 per audio hour, while the company's accuracy superlative lacks a benchmark in its announcement.

A person reviews an advanced transcript on a laptop screen, showing identified speakers, emotional cues like [laughter], and multilingual text.

Fish Audio said on September 28th that its speech-to-text model can mark speaker changes and add emotion and vocal-event labels, extending a product that turns audio into more structured transcripts. The company described the update in a post on X; its speech-to-text documentation identifies the model with those features as transcribe-1-pro.

https://x.com/fishaudio/status/2104617131920748857?s=46

poster=/api/storage/public-objects/tweet-videos/fish-audio-speaker-emotion-transcription-poster-f6f9ff31.jpg|Video from @FishAudio on X

The system labels speakers with numeric markers embedded in the transcript and adds bracketed cues such as laughter or an emotion. Those labels are machine inferences from the audio, not verified identities or objective measurements of a speaker's feelings. Fish Audio's documentation says speaker numbers identify voices only within a single recording; they do not identify the person across separate recordings. It also says the cues appear inline in the transcript rather than in a separate emotion field.

The distinction is practical for developers building transcript search, meeting notes, captions or voice products: the output can preserve some information about who spoke and how a passage sounded, rather than returning words alone. Fish Audio's API returns the transcript and can provide timestamped segments; its documentation notes that speaker and emotion markers are excluded from the timestamp alignment. Developers therefore need to handle the annotated transcript and timed speech as related, but not identical, outputs.

Fish Audio's post says the ASR model supports 83 languages and calls it the most accurate speech-to-text model available. The language figure and accuracy superlative are company claims; the post does not name a benchmark, comparison set or evaluation method for the accuracy claim. The product documentation names two models, transcribe-1 for general transcription and transcribe-1-pro for multi-speaker recordings with emotion and vocal-event cues. It does not substantiate the 83-language claim on the page reviewed here.

The API is available on a pay-as-you-go basis. Fish Audio's pricing documentation lists both transcription models at $0.36 per audio hour, with charges based on processed duration and rounded up to the nearest second. The company also says users can upload audio for transcription through its web app. The September post does not state whether this announcement changed pricing or availability, and Fish Audio's own site had already published a podcast-transcription guide on March 27th, 2026. The date therefore marks the company's post about the capability, not necessarily the first public availability of transcription at Fish Audio.

For co-founder and CEO Rissa Cao (@rissa_cao), the speech-to-text features fit a broader push to sell voice technology as infrastructure for products and businesses, alongside speech generation. In a July 27th company post announcing Fish Audio's seed round, Cao said her experience in voice AI at Amazon Alexa and Meta had shaped the company's focus on expressive delivery. Co-founder and chief scientist Shijia Liao, a former NVIDIA video researcher, started the underlying voice project after finding synthetic speech too monotonous for the VTuber content he watched, according to the company. The open-source project, Fish Speech, became the technical starting point for Fish Audio.

Fish Audio said in July that it raised $52 million in seed funding, led by Coreline Ventures and Capital Today, and reported $21 million in annual recurring revenue and more than 8 million users. Those figures come from the company and its fundraising coverage; they describe the business at the time of the July announcement, not independently audited results. The speech-to-text addition puts another product beside the voice-generation models that brought Fish Audio its early audience, while making the API useful for jobs where a plain transcript leaves out speaker turns and vocal events.

The broader commercial test is whether those annotations are dependable enough for customers to build workflows around them. Fish Audio's documentation describes how the labels are encoded and cautions that speaker markers are recording-specific. Its September post offers no comparative recognition results for this ASR model. For buyers weighing transcription services, the feature set is concrete; the accuracy claim remains a company assertion without a cited benchmark.

Reader comments

Conversation for this story loads after sign-in.