What it renders best
- A single well-lit person facing the camera. Identity lock across the whole take is the model’s superpower, and this is the photo that gives it the most to anchor to.
- Natural speech at 110–160 WPM. Every persona in the launch demo sits in this range; the pace accepts 80 to 220, but the top of the range sounds rushed and can smear lip closures.
- Voice and delivery driven by the scene prompt. The model generates audio and video together and conditions both on the prompt, so the prompt is where the voice is cast: register, grain, accent, breath, cadence, and dynamics.
Before & after
Casting the voice"A single person speaks directly to the camera. His voice is a low, salt-roughened baritone with real grain in it, chesty and open, worn down by wind and shouting across a deck. A rolling West Country lilt, consonants bitten off short, vowels stretched wide."A casting note: register, grain, accent, and cadence, describing a speaker rather than issuing an
instruction."in a rough, cheerful sailor's voice" or "make him sound deeper"An adjective label names a type without pinning a read, and two prompts written this way land on the
same generic voice."news anchor at a desk, professional tone": staging and delivery, which the scene prompt owns."change the shirt to red": identity and wardrobe are anchored to the photo. Text loses this fight
every time.The photo is the ceiling
There is no prompt that recovers a bad photo. A single well-lit, forward-facing photo anchors identity for the entire take and gives the strongest identity lock and the crispest lip-sync. A low-quality, occluded, or profile photo degrades everything downstream. A single-subject photo is strongly recommended.What each input controls
Treat these as three independent channels. A prompt that tries to reach across channels will be
ignored on that axis.
Writing scene prompts
- Describe the voice, staging, lighting, and tone, not the subject’s appearance.
- Cast the voice the way a director casts a part: register, grain, accent, breath, cadence, and dynamics. Describe the speaker in the third person rather than instructing the model.
- Delivery energy follows the scene prompt: an energetic prompt changes pacing and expression even with the same script.
- Keep the prompt under 800 characters, the limit
set_promptaccepts. - Send the prompt once and leave it alone. Resending
set_promptmid-conversation resets the voice, so a character that has been speaking will come back sounding like someone else. Pinset_seedto a fixed number for the same reason.