The character and scene were developed from a static image, then animated locally using ComfyUI and EchoMimic V3, with synthetic voice audio driving the initial facial performance and lip synchronization. The resulting footage was then refined with generative video processing to improve facial motion, expression, temporal consistency, and natural speech articulation.