Fish Audio Revolutionizes Voice Tech with Real-Time Cloning in 83+ Languages, Expands Enterprise Solutions

July 28, 2026
Fish Audio Revolutionizes Voice Tech with Real-Time Cloning in 83+ Languages, Expands Enterprise Solutions
  • Fish Audio is building a real-time, expressive voice platform that can clone a voice from a short clip in seconds, support 83+ languages, and offer more than 15,000 natural-language controls for emotion and pacing.

  • The platform provides text-to-speech, speech recognition, voice cloning, and long-form audio tools, hosting a massive library of over two million user-submitted voices and offering API access for developers and enterprises.

  • Its flagship model, S2.1 Pro, demonstrated strong performance in blind tests, with about two-thirds of listeners preferring its output over competitors.

  • The business mixes free/open access with paid enterprise plans, leveraging a rebuilt inference stack and FP8 GPU kernel library to keep the free tier affordable while serving commercial customers who need service-level guarantees.

  • Industry voices stress the importance of consent, transparent licensing, clear reporting and takedown workflows, and potential revenue sharing to build trust with creators.

  • Fish Audio operates in a competitive space with players like ElevenLabs and WellSaid, aiming to differentiate through fine-grained developer controls and cost-efficient model training.

  • The seed round features cross-border investors, signaling a trend of treating voice tech as an infrastructure layer and raising questions about governance, consent, and attribution for voice data.

  • Backers include Coreline Ventures, Capital Today, 359 Capital, and HF0 founder residency, highlighting strategic implications for a community-driven voice library and ongoing focus on consent and transparency.

  • The platform targets developers and regulated industries with HIPAA-compliant on-premises deployments, zero-data-retention policies, and plans to expand enterprise sales and API integrations with partners like Retell AI and LiveKit.

  • There is an ongoing governance dialogue over user-submitted voices, including a voice-ownership dispute process introduced recently and improvements in removal timelines after concerns about consent.

  • Open-weight models and a hosted API serve more than 8 million users and roughly $21 million in annual recurring revenue, supporting advances toward higher-quality models and broader enterprise sales.

  • The library now includes over 15,000 natural-language controls, serving both creative professionals and enterprises automating customer support and sales.

Summary based on 8 sources


Get a daily email with more Tech stories

More Stories