we prepped a full set of references before writing a single line of the video prompt, the chef, his daughter, the cart, the dish, the street, the interior, each one doing its own job so nothing drifted mid shot. the biggest surprise was around the video prompt itself, not the images. this is seedance 2.5 specifically, generating actual video, we made the reference images separately in chatgpt and going long and detailed there caused no problems at all. for the video though we first wrote something huge, choreographing every single beat down to the smallest detail, and it actually performed worse. not fully sure why, but it seems like over choreographing every little thing is what makes seedance choke, not what helps it. a shorter, more direct prompt worked better once the references were already doing the heavy lifting on identity and style.