- create an avatar from an image and save its id
- start a voice call and play the character
- publish your microphone, and send text with
say - interrupt the character, and change its voice or outfit mid-call
- end the call, and handle a call that ends on its own
Install the SDK
- npm
- pnpm
Understand the model
Three ideas cover most of the model:- The avatar and the call are separate. Create or attach an avatar, then start and end
calls with it. Ending a call keeps the avatar, so the next
start_callis fast. session_statereports the avatar and call state. It arrives on connect and on every change. Use it for the phase and controls; usetranscriptmessages for captions.- Model refusals and call failures arrive as
command_error. Branch on itscodeand useoriginto distinguish invalid input from a service failure. Handle SDK connection and upload errors separately.
Connect
Mint a token on your server, as Authentication describes, and pass it toconnect(). Add media elements to the page, then subscribe to messages and tracks before you
connect so you do not miss the first snapshot. jwtToken is the token your server returns.
Connect render, appendTranscript, and showError to your app’s UI.
webcam even for an audio call. The model sends its first
session_state as soon as the session connects. Wait for it before sending a command.
Browsers may block sound until the user interacts with the page. Start the call from a button
click. If autoplay is still blocked, the user can press play on the audio element.
Create or attach an avatar
On the first visit, upload an image and create the avatar. The snapshot moves topreparing_avatar,
then to avatar_ready. Save the avatar_id it reports.
command_error reports AVATAR_NOT_FOUND; clear the saved id and create a new avatar.
Start the call
Publish the microphone beforestart_call so it reaches the character when the call becomes
live. Speech before that phase may not reach the character.
list_voices reply:
{}. Use a voice value from the catalog in start_call.
Omit voice to use the default; do not send null.
While the phase is starting or warming_up, show a spinner. When the microphone reaches the
live call, the snapshot reports mic_forwarding: true. If the user blocks the microphone, start
the call anyway: say still reaches the character. To mute, set mic.enabled = false when mic
is set.
To let the character see the user, start the call with call_mode: "video", and publish the camera
to webcam the same way. call_mode is fixed for the call.
Talk and steer
With the microphone published, the user can just speak. Thetranscript messages give you both
sides of the conversation as captions. For typed chat, send say:
putOnJacket with a publicly fetchable
image URL, and keep the returned ID to remove that image later. image_id is a label you choose,
unique within the call; it does not need to be a UUID. Use a different ID for each additional image.
state.phase is live and state.control_ready is true. Enable the
controls from those two fields together. Outside a live call the commands are refused with
NOT_LIVE.
End the call
ending to ended. The avatar stays available for another call. When
the call ends, or start_call is refused, stop the microphone with mic?.stop() and call
reactor.unpublishTrack("mic"). A “Call again” button should publish the microphone again before
sending start_call.
A session bills while it is ready, including time between calls. Call reactor.disconnect() when
the user is done, or after a few idle minutes with no call.
A call can also end on its own. Handle it in render() from end_reason:
call_max_seconds and call_elapsed_seconds on the snapshot let you show a countdown before the
time limit ends the call.
Surface errors
Every refused command arrives ascommand_error, and the latest one also sits on
state.last_error.
start_call fails with UPSTREAM_CAPACITY, wait a few seconds before trying again. Log trace_id with every error so you can quote it when you report a failure.
What this tutorial leaves out
- Multiple clients. Several clients can join one session and steer the same character. See Multiple clients.
- Turn-taking and reply tuning. See the prompt guide.