Skip to main content
Use these commands to create or reuse an avatar, start a call, and change the character during the conversation.

Tracks

All four tracks are available when the session connects, so you can set up a player before a call. Character video and speech begin when the call is live. See Tracks for how to publish mic and webcam.

Session lifecycle

A session has one current avatar. You can start, end, and restart calls with it without reconnecting. Ending a call leaves the session open. Reactor sends a session_state message when you connect and whenever the avatar or call changes. Its phase field reports the current step: If avatar creation fails, the phase returns to idle. If a call fails, the phase becomes failed. Two limits can end a call without end_call. call_max_seconds is the maximum call length; its timer begins when the call starts. Reaching it sets end_reason to max_duration. Two hours without input sets end_reason to idle_timeout.

Commands

Send commands with reactor.sendCommand() in JavaScript or reactor.send_command() in Python. A command with a named reply returns that message. Other commands return no value; watch session_state for their result. A refused command sends command_error without changing the session. See Command errors for the codes.
Omit an optional field you do not set. Do not send it as null. A command that carries an explicit null is dropped whole, without a command_error, so send { persona }, not { persona, voice: null }.

create_avatar

Create an avatar from one image of a person. The new avatar becomes the current avatar for this session. Give exactly one of image or image_url. An uploaded image can be PNG, JPG, WebP, or HEIC, under 20 MB and at most 50 megapixels. The model applies the image’s EXIF orientation, so a phone photo arrives upright. While the image is processed, session_state.phase is preparing_avatar. When the avatar is ready, the phase becomes avatar_ready and session_state includes its avatar_id. Save the ID to use the avatar in another session.

attach_avatar

Use an avatar you created earlier by passing its avatar_id. Avatars are kept for 90 days. The model checks the ID before using the avatar. On success, session_state.phase becomes avatar_ready.

list_voices

Send {} at any time. The voices reply lists built-in voices and names the default voice. Each system entry has voice, description, and accent. Pass an entry’s voice to start_call or update_call.

start_call

Start a live call with the current avatar. Create or attach an avatar first, then wait for session_state.phase to become avatar_ready. Only persona is required.
string
required
Who the character is and how it behaves, up to 50,000 characters.
string
A voice from list_voices. Omit it to use the model’s default voice.
string
What the character says or does first, before you speak, up to 200 characters.
string
default:"English"
Language the character speaks and replies in, up to 40 characters.
string
default:"audio"
audio forwards your mic track. video also forwards your webcam track. This setting is fixed for the call.
boolean
default:"true"
Send a transcript message for what the caller and character say.
boolean
default:"false"
Expand a short persona into a fuller description before the call starts. Use this when you have only a brief character description.
object
Turn-taking settings. You can change them later with update_call.
object
Reply generation settings. You can change them later with update_call.
call_mode controls whether your camera reaches the character. The character’s generated video still arrives on main_video in both modes.

Turn-taking: vad

VAD means voice activity detection. These settings control when caller speech starts or ends a turn. The default server mode filters brief acknowledgments and background noise. Use semantic when the caller should be able to interrupt the character as soon as they begin speaking.
string
default:"server"
server filters back-channel sounds and background noise. semantic lets the caller interrupt the character as soon as they start speaking.
number
default:"0.5"
With server, controls how much noise is filtered, from 0 to 1. Raise it in a noisy room; lower it if quiet speech is missed.
integer
default:"400"
With server, how long the caller must be silent before the character answers, from 200 to 6000 milliseconds. Raise it to allow longer pauses; lower it for faster replies.
See the turn-taking examples for common settings.

Reply generation: llm

These settings shape each reply. A token is a piece of text. You can change these settings during a call with update_call; the new values apply on the next turn.
integer
default:"50"
The maximum length of one reply in tokens. Raise it for longer answers; lower it for brief exchanges.
number
Controls variation, from 0 up to but not including 2. Lower values make replies more consistent; higher values make them more varied.
number
Nucleus sampling, greater than 0 and at most 1. Lower values narrow the next-token choices, making replies less varied.
integer
Top-k sampling, from 0 to 100. A lower positive value limits how many next-token choices the model considers.
number
Discourages repeated words and phrases. Raise it when replies repeat themselves; the value must be greater than 0.
number
Encourages new topics instead of returning to ones already covered. Raise it for more topic variety, from 0 to 2.
integer
Controls randomness. Use a fixed value to make a demo more repeatable; -1 picks a random seed.
The phase moves through starting and warming_up to live. Once live, main_video and main_audio carry the character, and your mic reaches it. If the call fails, the phase becomes failed; inspect command_error or session_state.last_error for the reason.

say

Send text for the character to answer, as if you had spoken it. There is no reply. The answer arrives on main_video and main_audio, and as a transcript when transcripts are on.

interrupt

Send {} to stop the character mid-sentence, as speaking over it would. There is no reply. To stop it and give a new instruction, send interrupt followed by say; interrupt takes no text.

update_call

Change the live call without restarting it. Give at least one field. The model applies all supplied fields in one update. If it rejects any field, none of the changes take effect. The call_updated reply lists the settings that changed.

set_reference_images

Give the character one to three reference images to pick up live: an object to hold, an outfit to wear, or a background. Each entry has these fields:
string
required
A publicly fetchable image. Reference images are URL-only; uploads are not accepted here.
string
required
Your own stable id for the image, 1 to 128 characters, unique in the call. Sending an id that is already in effect replaces that image.
string
object, garment, or background. Inferred when omitted.
string
One sentence, up to 200 characters, that describes the change.
The change does not interrupt speech. The reference_images_applied reply lists every image ID now in effect.

clear_reference_images

Remove reference images from the live call. Pass the image_id values you supplied with set_reference_images. Omit image_ids to undo the most recent set_reference_images command. The reference_images_applied reply lists the IDs still in effect. This example removes the jacket from the preceding set_reference_images example:

end_call

Send {} to end the call. The call_ended reply arrives once the call is released. Its duration_seconds is the time the call was live, or 0 if it never became live. The phase moves through ending to ended. The avatar remains available, so start_call works again.

get_state

Send {} at any time to get the current session_state. Reactor also broadcasts this message when the avatar or call changes.

Messages

Every message is broadcast to all connected clients, except command replies, which go only to the client that sent the command. For transcript, speaker is user or character, and final is true when the text is settled. Listen for messages before connecting so you receive the first session_state:

session_state

The authoritative snapshot. Drive your UI from this message alone. A client that reconnects mid session gets a full snapshot on connect. Only phase is always present; the other fields are null or false when they do not apply. end_reason is one of: Any value other than ended_by_client also leaves a last_error that explains it.

command_error

Emitted when a command is refused or a call fails. A refused command does not perform the requested action; a call failure moves the phase to failed. The payload is also stored as session_state.last_error. origin is one of request (your input was invalid), state (the command is not allowed in this phase), upstream (the generation service refused or failed), or platform (a failure on Reactor’s side). The generation service can also return its own codes, which pass through unchanged. Examples are LIVE_CONN_INIT_FAILED and codes that start with PROMPT_OP_.

Upload reference

create_avatar.image takes a reference returned by the Reactor upload protocol. All fields are required.

Multiple clients

Several clients can join the same session. session_state and command_error go to all of them; a command reply goes only to the client that sent the command. A session has one call, so every client steers the same character. A client that disconnects does not end the call. Every new session starts in idle, without an avatar.