Fish Audio is an artificial intelligence company that develops speech and audio technologies, including text-to-speech, voice cloning and speech-to-text systems
Fish Audio is an artificial intelligence company that develops speech and audio technologies, including text-to-speech, voice cloning and speech-to-text systems. The company operates the Fish Audio platform and develops speech models, including Fish-Speech and the S2 series.[1][2]
| Type | Private |
|---|---|
| Industry | Artificial intelligence |
| Founded | 2025 |
| Founders | Shijia Liao RissaCao |
| Products | Text-to-speech, voice cloning, speech-to-text |
| Parent | Hanabi AI Inc. |
| Website | https://fish.audio/ |
Fish Audio originated as a project developed by Shijia Liao, a former video researcher at Nvidia. The project focused on synthetic speech and was initially developed using a single graphics processing unit. The resulting Fish-Speech project was released as an open-source text-to-speech system. In 2024, Liao and other researchers published a technical paper describing Fish-Speech, a multilingual text-to-speech framework based on a large language model architecture. Rissa Cao joined Liao as a co-founder and became chief executive officer of Fish Audio. The company subsequently developed commercial services based on its speech-generation technology.[3][4]
Fish Audio develops artificial intelligence systems for speech synthesis, voice cloning and related audio applications. Its technology includes text-to-speech generation, voice cloning and speech-to-text. In March 2026, the company released the S2 model and published a technical report describing it as a text-to-speech system supporting multi-speaker and multi-turn generation and instruction-based control using natural-language descriptions. The release included model weights, fine-tuning code and an inference engine. The S2 model uses a Dual-Autoregressive architecture and generates speech based on textual instructions describing characteristics of the output. Fish Audio's models have been released in hosted and downloadable forms. The S2 model is distributed under a license that places restrictions on commercial use.[5][6]
Fish Audio operates a web-based platform providing speech-generation and audio-processing services. Its services include text-to-speech, voice cloning, speech-to-text, voice changing, audio separation and audio translation. The company also provides application programming interfaces for integrating its speech-generation systems into other applications. In July 2026, Fish Audio announced a $52 million seed funding round led by Coreline Ventures and Capital Today. The round included participation from 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners and Alphalist Partners.[7][8]
Fish Audio's research work includes the Fish-Speech project, an open-source text-to-speech system released through public software repositories. The Fish-Speech research paper was published in 2024 and described a multilingual text-to-speech architecture using large language models for linguistic feature processing. The project included implementations and model resources for researchers and developers.[9][10]
Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.