Skip to main content
Vidu S2-Avatar has no scene prompt. You shape the character with five inputs: the image you build the avatar from, the persona and greeting you send with start_call, the text you send with say, and the reference images you give it during the call. This page covers each one.

Image

For the image input, use one image of one person. Use a full-body or half-body shot with the face clearly visible. The avatar keeps this look for every call, so pick an image that matches the character you plan to write.
  • Use one person per image. Other people in the frame can confuse the avatar.
  • Use a front-facing or three-quarter view, with even light on the face.
  • A phone photo is fine. The model applies its EXIF orientation, so it arrives upright.

Persona

persona is required on start_call. It says who the character is and how it behaves. It can be up to 50,000 characters, but a short, specific persona usually works better than a long one. A good persona covers four things: Compare these two personas:
You are a helpful assistant.
You are Tina, a guide at a sea-life museum. You are warm and curious, and you love octopuses. Answer in one or two short sentences, then ask the user a question. If asked about tickets or prices, say the front desk can help.
The first persona gives the character no voice of its own, so it answers like a generic assistant. The second gives it a name, a setting, a manner, and a clear limit on length.
Keep answers short. The character speaks every word it generates, and a long answer is a long wait before the user can reply. The llm.max_tokens setting, 50 by default, also caps each reply. See Tune replies.
Set persona_enhance: true to let the model expand a short persona into a fuller one before the call starts. Use it when you have only a line or two. Leave it off when you wrote the persona with care and want it used as is.

Greeting

greeting is optional, up to 200 characters. It says what the character says or does first, before you speak. Without it, the character waits for you. Treat the greeting as a short instruction for the opening. For example:
The character generates the opening from this instruction. Test the result in a call if specific words or gestures matter to your app.

Send text with say

say sends text as if you had spoken it, up to 2,000 characters. Use it for typed chat, for a scripted demo, or to test a persona without a microphone. The character answers on the tracks, the same as it answers speech. say is input from the user’s side of the conversation, not a line for the character. To change what the character says or how it behaves, change the persona with update_call.

Choose a language and a voice

start_call accepts an optional language field. If you omit it, the model follows the conversation. The examples in this guide use English. Call list_voices to see the voices you can use. Each built-in voice has a description and an accent. Pick one that fits the image and the persona, and pass its voice to start_call. You can change it mid-call with update_call; the new voice starts after the current sentence.

Tune turn-taking

vad controls when the character decides you have finished speaking. The default server mode filters back-channel sounds, such as “mm-hm”, and background noise. The threshold and silence settings above apply to this mode. With semantic, the character stops when the user starts speaking. To stop it from your app, send interrupt.

Tune replies

llm controls how the character generates each reply.
  • max_tokens, 50 by default, caps the length of one reply. Raise it for a storyteller; keep it low for quick back-and-forth.
  • temperature sets how varied the replies are. Lower values give steadier, more predictable answers.
  • frequency_penalty adjusts repetition. 1 is neutral; lower values encourage repeated wording, and higher values discourage it.
  • seed sets the random seed for reply generation. -1 picks a random seed.
Change vad and llm mid-call with update_call. They apply on the next turn.

Change the look mid-call

set_reference_images gives the character up to three images to pick up while it talks. Use kind to tell the model what each image represents: If you omit kind, the model looks at the image and guesses whether it is an object, a garment, or a background. Set kind when an image contains more than one possible target, such as a jacket against a distinctive background. text is optional, one sentence of up to 200 characters. Use it to say what happens, in the same plain way you would describe the change to a person. Images must be publicly fetchable URLs. Use an image with a plain background for an object or a garment, so the character picks up the item and not its surroundings. Give each new reference image its own image_id and keep the IDs you use. To remove a specific image, send its ID with clear_reference_images. Omit image_ids to undo the most recent set_reference_images command.

Handle content refusals

Vidu S2-Avatar is moderated. A refused request arrives as command_error with code: "CONTENT_POLICY". A call stopped for its content ends with end_reason: "content_policy", and last_error.reason says why. Show that reason to the user rather than retrying the same text.

See also

  • Schema: every command, parameter, and message.
  • Tutorial: build a browser call from start to finish.
  • Content moderation: how Reactor moderates the models it serves.