When two people talk in one frame, identities blur, both mouths move and lip sync fails.
Example scene
An office. Zhou Yao slams a file on her manager Lin Xia's desk and they trade 4 lines. 15 seconds total, split into one two-shot and four single close-ups.
Steps
Make the two look different: Lin Xia with short hair and a black blazer, Zhou Yao with long curls and a cream sweater.
Set the axis: Lin Xia always frame left looking right, Zhou Yao always frame right looking left.
Generate one two-shot with a single action, slamming the file, and no dialogue.
Generate a face-on single close-up for each line, with eyelines matching the axis.
Lock all 4 voice lines first, then generate lip sync for each close-up to the audio length.
Edit as two-shot → speaker → reaction → speaker, cutting to the listener just before each line ends.
Prompts
Two-shot
Open-plan office. Frame left: Lin Xia, short hair, black blazer, seated behind a desk. Frame right: Zhou Yao, long curly hair, cream sweater, standing in front of the desk. Zhou Yao slams a file onto the desk; Lin Xia looks up at her. A desk between them, medium shot, side angle, vertical 9:16. Only these two people in frame.
Single close-up with dialogue
Chest-up close-up of Zhou Yao, facing camera, looking slightly off frame left, saying with suppressed anger: “You didn't even read this report.” Natural mouth movement, pauses matching the audio, head leaning in only slightly. Keep her long curls, cream sweater and background unchanged. Only her in frame.
Negative
extra people, both people talking at once, merged faces, swapped clothing, talking in profile, mouth covered
✕ Failed take
One 10-second two-shot with both characters trading 4 lines: their hair gradually converges, Zhou Yao's mouth moves while Lin Xia speaks, and at second 7 they swap jacket colours.
✓ Passing take
Across 5 shots, only one mouth moves per line. Pass criteria: the axis is never crossed and left/right positions hold throughout; every line's mouth open and close match the audio; with the sound off you can still tell who's talking.
Before/after clips are in production. Follow the steps and prompts to reproduce the result.