INPUTLocked dialogue, voice profiles, and picture
DELIVERABLEDialogue, ambience, effects, music, and mix
Tools: MiniMax Speech, CapCut, and licensed sound libraries
Follow the workflow
Record perceived age, pitch, pace, breath, emotional range, and exclusions per character.
Generate and audition dialogue line by line before changing picture timing.
Lay continuous room tone under every scene.
Place action effects exactly on events and use music only for mood and rhythm.
Audition on a phone speaker and reduce anything masking dialogue.
Reusable template
Character: [age/personality/current state]. Voice: [pitch, texture, pace, breath]. Emotion: [type and intensity]. Line: “[text]”. Performance: [pause, stress, ending]. Avoid announcer delivery, exaggeration, and age drift.
Keep invariant fields fixed and replace only project variables. Change one variable per test.
Quality gates
- A character sounds consistent across shots
- Dialogue feels natural and fits the action
- Ambience bridges edits
- Music never masks speech and every asset is licensed
Troubleshooting:
Add listener, subtext, and pauses to fix narration-like delivery. Reduce emotional intensity when acting is excessive. Use room tone to hide edit seams.
Add listener, subtext, and pauses to fix narration-like delivery. Reduce emotional intensity when acting is excessive. Use room tone to hide edit seams.
SPECIALIST GUIDES
Additional practical guidance
Lock final dialogue audio before lip sync. Prefer frontal, unobstructed faces and restrained head motion.
Give one line to one clearly identified speaker; split multi-speaker scenes into singles and reaction shots when needed.
Check mouth opening, closure, consonant attacks, pauses, and line endings—not only a few middle frames.
Build four layers in order: dialogue, continuous ambience, timed effects, then music.
Chapter completeDialogue, ambience, effects, music, and mix. Save the approved versions, prompts, and revision notes before handoff.
KEEP LEARNING