Tracks
All four tracks are available when the session connects, so you can set up a player before a call.
Character video and speech begin when the call is live. See Tracks for how to
publish
mic and webcam.
Session lifecycle
A session has one current avatar. You can start, end, and restart calls with it without reconnecting. Ending a call leaves the session open. Reactor sends asession_state message when you connect and whenever the avatar or call changes.
Its phase field reports the current step:
If avatar creation fails, the phase returns to
idle. If a call fails, the phase becomes failed.
Two limits can end a call without end_call. call_max_seconds is the maximum call length; its
timer begins when the call starts. Reaching it sets end_reason to max_duration. Two hours
without input sets end_reason to idle_timeout.
Commands
Send commands withreactor.sendCommand() in JavaScript or reactor.send_command() in Python. A
command with a named reply returns that message. Other commands return no value; watch
session_state for their result. A refused command sends command_error without changing the
session. See Command errors for the codes.
create_avatar
Create an avatar from one image of a person. The new avatar becomes the current avatar for this
session. Give exactly one of image or image_url.
An uploaded image can be PNG, JPG, WebP, or HEIC, under 20 MB and at most 50 megapixels. The model
applies the image’s EXIF orientation, so a phone photo arrives upright.
While the image is processed,
session_state.phase is preparing_avatar. When the avatar is ready,
the phase becomes avatar_ready and session_state includes its avatar_id. Save the ID to use
the avatar in another session.
attach_avatar
Use an avatar you created earlier by passing its avatar_id. Avatars are kept for 90 days.
The model checks the ID before using the avatar. On success,
session_state.phase becomes
avatar_ready.
list_voices
Send {} at any time. The voices reply lists built-in voices and names the default voice. Each
system entry has voice, description, and accent. Pass an entry’s voice to start_call or
update_call.
start_call
Start a live call with the current avatar. Create or attach an avatar first, then wait for
session_state.phase to become avatar_ready. Only persona is required.
string
required
Who the character is and how it behaves, up to 50,000 characters.
string
A
voice from list_voices. Omit it to use the model’s default voice.string
What the character says or does first, before you speak, up to 200 characters.
string
default:"English"
Language the character speaks and replies in, up to 40 characters.
string
default:"audio"
audio forwards your mic track. video also forwards your webcam track. This setting is
fixed for the call.boolean
default:"true"
Send a
transcript message for what the caller and character say.boolean
default:"false"
Expand a short persona into a fuller description before the call starts. Use this when you have
only a brief character description.
object
Turn-taking settings. You can change them later with
update_call.object
Reply generation settings. You can change them later with
update_call.call_mode controls whether your camera reaches the character. The character’s generated video
still arrives on main_video in both modes.
Turn-taking: vad
VAD means voice activity detection. These settings control when caller speech starts or ends a turn.
The default server mode filters brief acknowledgments and background noise. Use semantic when
the caller should be able to interrupt the character as soon as they begin speaking.
string
default:"server"
server filters back-channel sounds and background noise. semantic lets the caller interrupt
the character as soon as they start speaking.number
default:"0.5"
With
server, controls how much noise is filtered, from 0 to 1. Raise it in a noisy room; lower
it if quiet speech is missed.integer
default:"400"
With
server, how long the caller must be silent before the character answers, from 200 to 6000
milliseconds. Raise it to allow longer pauses; lower it for faster replies.Reply generation: llm
These settings shape each reply. A token is a piece of text. You can change these settings during a
call with update_call; the new values apply on the next turn.
integer
default:"50"
The maximum length of one reply in tokens. Raise it for longer answers; lower it for brief
exchanges.
number
Controls variation, from 0 up to but not including 2. Lower values make replies more consistent;
higher values make them more varied.
number
Nucleus sampling, greater than 0 and at most 1. Lower values narrow the next-token choices, making
replies less varied.
integer
Top-k sampling, from 0 to 100. A lower positive value limits how many next-token choices the model
considers.
number
Discourages repeated words and phrases. Raise it when replies repeat themselves; the value must be
greater than 0.
number
Encourages new topics instead of returning to ones already covered. Raise it for more topic
variety, from 0 to 2.
integer
Controls randomness. Use a fixed value to make a demo more repeatable;
-1 picks a random seed.starting and warming_up to live. Once live, main_video and
main_audio carry the character, and your mic reaches it. If the call fails, the phase becomes
failed; inspect command_error or session_state.last_error for the reason.
say
Send text for the character to answer, as if you had spoken it.
There is no reply. The answer arrives on
main_video and main_audio, and as a transcript when
transcripts are on.
interrupt
Send {} to stop the character mid-sentence, as speaking over it would. There is no reply. To stop
it and give a new instruction, send interrupt followed by say; interrupt takes no text.
update_call
Change the live call without restarting it. Give at least one field.
The model applies all supplied fields in one update. If it rejects any field, none of the changes
take effect. The
call_updated reply lists the settings that changed.
set_reference_images
Give the character one to three reference images to pick up live: an object to hold, an outfit to
wear, or a background.
Each entry has these fields:
string
required
A publicly fetchable image. Reference images are URL-only; uploads are not accepted here.
string
required
Your own stable id for the image, 1 to 128 characters, unique in the call. Sending an id that is
already in effect replaces that image.
string
object, garment, or background. Inferred when omitted.string
One sentence, up to 200 characters, that describes the change.
reference_images_applied reply lists every image ID now
in effect.
clear_reference_images
Remove reference images from the live call. Pass the image_id values you supplied with
set_reference_images. Omit image_ids to undo the most recent set_reference_images command.
The
reference_images_applied reply lists the IDs still in effect. This example removes the jacket
from the preceding set_reference_images example:
end_call
Send {} to end the call. The call_ended reply arrives once the call is released. Its
duration_seconds is the time the call was live, or 0 if it never became live. The phase moves
through ending to ended. The avatar remains available, so start_call works again.
get_state
Send {} at any time to get the current session_state. Reactor also broadcasts this message when
the avatar or call changes.
Messages
Every message is broadcast to all connected clients, except command replies, which go only to the client that sent the command.
For
transcript, speaker is user or character, and final is true when the text is
settled. Listen for messages before connecting so you receive the first session_state:
session_state
The authoritative snapshot. Drive your UI from this message alone. A client that reconnects mid
session gets a full snapshot on connect. Only phase is always present; the other fields are null
or false when they do not apply.
end_reason is one of:
Any value other than
ended_by_client also leaves a last_error that explains it.
command_error
Emitted when a command is refused or a call fails. A refused command does not perform the requested
action; a call failure moves the phase to failed. The payload is also stored as
session_state.last_error.
origin is one of request (your input was invalid), state (the command is not allowed in this
phase), upstream (the generation service refused or failed), or platform (a failure on Reactor’s
side).
The generation service can also return its own codes, which pass through unchanged. Examples are
LIVE_CONN_INIT_FAILED and codes that start with PROMPT_OP_.
Upload reference
create_avatar.image takes a reference returned by the Reactor upload protocol. All fields are
required.
Multiple clients
Several clients can join the same session.session_state and command_error go to all of them; a
command reply goes only to the client that sent the command. A session has one call, so every client
steers the same character. A client that disconnects does not end the call. Every new session starts
in idle, without an avatar.