Fish Audio raises $50M seed to build AI voice models
Fish Audio raises 50m in a seed round led by Coreline Ventures and Capital Today to expand AI voice models for creators and enterprises. The Palo Alto startup already serves more than 8 million users and posts $21 million in annual recurring revenue after launching last year.
Key Takeaways
- Fish Audio closed a $50 million seed round led by Coreline Ventures and Capital Today, with several other firms participating.
- More than 8 million people use its open-source or hosted models, and the company reports $21 million in annual recurring revenue.
- The startup offers expressive voice generation for creators and steerable voices for enterprise support and sales use cases.
- It has automated voice take-downs after consent complaints and plans audio understanding and speech-to-speech models next.
The raise, reported by TechCrunch, aims to fund more advanced models as Fish Audio moves deeper into enterprise customers. Investors in the round also included 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.
For readers tracking Future Tech & AI Wonders, the deal stands out because Fish Audio grew quickly without raising earlier, then sought capital as enterprise demand and model ambitions scaled.
What does Fish Audio actually build?
Fish Audio is building AI voice models meant to be more expressive for creative work and more steerable for enterprises that want to automate customer support and sales operations. The company says its library includes more than 15,000 natural language controls.
The project began when former NVIDIA researcher Shijia Liao trained a voice generation model on a single GPU and open-sourced it. The Fish Speech repository on GitHub now has more than 31,000 stars and is used by indie developers, game designers, and creators.
Over the past year, Fish Audio launched five models: four speech generation models and one speech-to-text model. It open-sourced three speech generation models, while its latest S2.1 Pro model is available only through a paid API. Creators and teams can buy monthly plans with generation minutes and voice cloning, and enterprises can use a dedicated API and platform. Organizations such as HeyGen, Sanas, and Plaud are already customers, according to the company.
Why did Fish Audio raise capital now?
CEO and co-founder Rissa Cao told TechCrunch the company ran efficiently when it was mainly an open-source project with creator plans and did not need outside money. It sought funding to develop more advanced models and serve enterprises as investor interest grew.
Looking ahead, Fish Audio plans to release an audio understanding model this year and is also building a speech-to-speech model. The speech generation market remains crowded, with rivals including ElevenLabs, WellSaid, Cartesia, Speechify, Async, and Krisp.
Rico Mallozzi, a partner at 359 Capital, said Fish Audio’s fine-grained developer controls and cost-efficient training help it compete with larger AI labs, calling its progress on more human-like voices technically impressive given the team’s size.
How is Fish Audio handling voice consent concerns?
One way the startup builds its voice library is by asking users to submit voices for model training and compensating them if those voices are used. That approach drew scrutiny months ago when some creators alleged their voices were uploaded without consent.
Fish Audio already had a DMCA take-down process, but removals were slow. Cao said the company has now automated take-downs: creators can submit a short voice sample or a contract to prove ownership, and a voice can be removed in less than three minutes. The system still depends on rights holders discovering misuse and filing a claim.
Oskue Honda, a partner at Coreline Ventures, argued a community-driven model only becomes durable if creators trust the platform, calling for consent, transparency, attribution, verified ownership, clearer licensing, and eventual revenue sharing when voices are licensed commercially.