FishAudio — S2.1 Pro Launch Video by Yunis MikayilovFishAudio — S2.1 Pro Launch Video by Yunis Mikayilov

FishAudio — S2.1 Pro Launch Video

Yunis Mikayilov

Yunis Mikayilov

Fish Audio builds open-source text-to-speech. S2.1 Pro is their production model: 83 languages, voice cloning, roughly 70ms to first audio, and natural-language control over emotion and delivery written inline with the text. They launched it with free unlimited API access, into a market where every competitor's best voice sits behind a paywall. The news was the model. The problem was showing it.
Most product films are built picture-first, with sound in support. This one inverts. The product is the sound, which meant every visual decision had to stay out of the audio's way and still hold attention for the length of the piece. We built the motion to track the voice rather than compete with it: waveforms, latency, language switches, and emotional tags treated as events the audio drives rather than decoration laid over the top. The model demonstrates itself. The animation is what makes that legible.
Film end to end: concept, design, and animation. It works as a launch asset and a demo at once - you hear what the model does while you watch what it means.
Like this project

Posted Jul 14, 2026

FishAudio — S2.1 Pro Launch Video. Product-led animation for the new model's launch.