Fish Audio Raises $52 Million Seed Round After Turning a Bedroom GPU Project Into One of Voice AI's Fastest-Growing Companies
Business📅 July 28, 2026👤 FreeReadText Team

Fish Audio Raises $52 Million Seed Round After Turning a Bedroom GPU Project Into One of Voice AI's Fastest-Growing Companies

Palo Alto-based Fish Audio raises a $52 million seed round co-led by Coreline Ventures and Capital Today, scaling from a solo developer's open-source side project into an 8-million-user platform with $21 million in annual recurring revenue — all within its first year.

On July 28, 2026, Fish Audio — a Palo Alto-based AI voice startup — announced a $52 million seed funding round co-led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, HF0, and 645 Ventures. The round is one of the largest seed-stage raises in voice AI history and caps a remarkable trajectory: the company began as a solo developer's passion project on a single gaming GPU and has grown to over 8 million users and $21 million in annual recurring revenue within its first year of operation.

Fish Audio was founded by Shijia Liao, a former NVIDIA video researcher, who started the project out of frustration with the flat, robotic quality of available synthetic voices. Liao — a lifelong VTuber and anime fan — trained an initial voice generation model on a single gaming GPU in his bedroom and open-sourced it as Fish Speech on GitHub, where it rapidly amassed more than 31,000 stars and became one of the most popular voice AI projects on the platform. Liao then partnered with CEO Rissa Cao to commercialize the technology. The company has since launched five models — four speech generation and one speech-to-text — with three remaining open-source and the flagship S2.1 Pro available exclusively through a paid API. The model supports 83+ languages and 15,000+ natural language controls for word-level emotion and pacing, with voice cloning from just 5 seconds of audio in approximately 15 seconds.

The open-source-to-commercial trajectory mirrors strategies employed by Meta and Mistral in the text LLM space, but Fish Audio is among the first voice AI companies to execute it at scale. Enterprise customers including HeyGen (AI avatars), Sanas (accent technology), LiveKit (real-time voice infrastructure), and Retell (phone agents) have adopted Fish Audio's models for production workloads, drawn by pricing of approximately $15 per million characters — significantly below ElevenLabs' enterprise tier of $60–$165 per million. The company also offers on-premises deployment, zero-data-retention policies, and HIPAA-compliant configurations for regulated industries, positioning it as a credible enterprise alternative to closed-source voice APIs.

The seed capital will fund development of voice-native large language models, speech-to-speech architectures, and a full audio-native stack that goes beyond TTS into reasoning-capable voice agents. Fish Audio also plans to build an enterprise sales team and deepen integrations with partners including LiveKit and Retell. The round reflects broader investor conviction in voice AI as an independent category: more than $7 billion flowed into voice AI startups in Q1 2026 alone, and Fish Audio's combination of open-source community gravity, enterprise revenue, and technical velocity positions it as one of the most closely watched companies in the space. As one investor noted, the company's challenge now is managing the tension between its open-source roots and the commercial demands of its fast-growing enterprise business.

Fish AudioSeed FundingOpen Source TTSVoice CloningCoreline VenturesCapital Today$52M

来源

← Back to News