Most product films are built picture-first, with sound in support. This one inverts. The product is the sound, which meant every visual decision had to stay out of the audio's way and still hold attention for the length of the piece. We built the motion to track the voice rather than compete with it: waveforms, latency, language switches, and emotional tags treated as events the audio drives rather than decoration laid over the top. The model demonstrates itself. The animation is what makes that legible.