I built a worked run of a long-talk-to-short-clips job on placeholder inputs. A placeholder transcript of a 45:00 talking-head recording (16 excerpted segments) is split into sentences and word timings, and every segment is scored with readable rules: a hook in the first sentence, a penalty if it leans on something said earlier, a cost per filler word, a point when it refers to something on screen. The best hook in each of 3 topics plus its best supporting segment becomes a clip; fillers are cut as jump cuts, punch-ins keep every shot at or under 3.9 s, and the demo recording or a screenshot cuts in only over the sentences that point at them (6 inserts). Clips run 19, 23 and 17 s, the first word lands at 0.0 s with no logo sting, and captions of at most 3 words cover the whole clip. It renders 6 MP4s, each clip at 1080x1920 (9:16) and 1080x1350 (4:5), with an edit decision list per clip. The speaker is a drawn placeholder and the voice is a placeholder tone track timed to the caption words; with real footage the same edit runs on the speaker's own audio and video.