To keep both characters consistent across every shot, I created detailed reference sheets for each physical description, wardrobe, posture, and a performance-progression ladder mapping how her expression shifts from shot to shot then locked those as FacePass references before batch generating. The trickiest element was a screen graphic (the extortion demand) that needed to stay perfectly legible across multiple takes, so I composited it in as a static asset rather than letting the model regenerate small text each time. For sound, I built a looping piano motif with a deliberate key-change timed to land exactly on the reveal, so the twist registers in the ear a half-beat before the eye catches up and a single unresolved low tone carries through to the final cut to black, so the ending stays unsettled instead of neatly resolved.