Why it happens
- The location word implies crowds, such as “street”, “mall” or “restaurant”.
- The first frame has a blurry figure, mirror or glass reflection that the model animates into a person.
- Long clips and big camera moves let the model fill newly revealed areas with people.
How to fix it
State the headcount in both keyframe and video prompts: only one person in frame.
Add qualifiers to the location, such as “completely empty” or “no passers-by late at night”.
Inspect the first-frame background and remove figure-like shapes and mirror or glass reflections.
Shorten the clip or reduce camera movement so the model has less new area to fill.
Add to the negative prompt: extra people, passers-by, crowd, people in reflections.
Prevent it next time:State the headcount in every prompt and leave nothing person-shaped in the first-frame background.
Worked examples