@reactor-models/sana-streaming SDK. By the end you’ll know how to edit your webcam feed in real time, edit an
uploaded clip of your choice, steer the edit mid-stream, snap clips, and surface model errors.
Installation and setup
Get the example running before reading further. Every section below points back at code in the example repo. You will need:- Node.js 18+.
- pnpm (the example’s lockfile is pnpm’s;
npmoryarnwork too). - A Reactor API key (starts with
rk_). - Familiarity with the Next.js App Router.
1
Clone the example
The example lives alongside our other reference apps in
reactor-team/js-sdk under examples/.2
Add your API key
Your
rk_… key must never reach the browser; the example reads it server-side and mints a
short-lived JWT for the client (the standard broker pattern); for now, drop
the key into .env.local:3
Install dependencies and start the dev server
http://localhost:3000, click Connect, allow camera access, pick a preset prompt (or
type your own edit), and press Start live.How SANA-Streaming works
Building with SANA-Streaming is different from Reactor’s other models. Helios, LingBot, and LongLive-2.0 generate video from a prompt; SANA-Streaming edits the video you bring. You open a long-lived connection, give the model a source (your webcam or an uploaded clip) and an edit instruction, and it streams back transformed frames in 24-frame chunks, one every ~1-1.5s. Re-prompt at any time and the new edit lands at the next chunk boundary, with no re-render and no break in the stream. Opening the connection isn’t instant. Reactor provisions a GPU for your session, so the client moves through four states before media starts flowing:waiting state is when the GPU is being assigned, which takes a few seconds. Once the status
reaches ready, commands take effect and the
session lifecycle begins in its idle
state. StatusBadge.tsx surfaces every connection state with a label and a Connect / Disconnect
toggle. See Sessions for the full breakdown.
Two properties of the API are worth internalizing before you read on:
- Commands are asynchronous; messages are the source of truth. Calling
setVideodoesn’t mean the model has a source yet; it confirms with avideo_acceptedmessage and astatesnapshot whosehas_videoflips totrue. - Errors arrive out-of-band. A broken precondition like
startwith no source surfaces later as acommand_errormessage (useSanaStreamingCommandError), not a thrown exception.
@reactor-models/sana-streaming SDK, which wraps the base SDK with
the model name and tracks baked in. SanaStreamingApp.tsx mounts
<SanaStreamingProvider getJwt={fetchToken}>; components read status and call one typed method per
command off useSanaStreaming() (setPrompt({ prompt }) rather than
sendCommand("set_prompt", { prompt })), and subscribe to messages with per-message hooks like
useSanaStreamingState. The wire shapes behind every method and message are in the
schema.
Clip recording is the one exception. It is model-agnostic and not re-exported by the typed
package, so
SnapClip.tsx uses the base @reactor-team/js-sdk
directly (useReactor, <ClipPlayer>). Drop it into any model example unchanged.The model is the source of truth
The browser sends commands and renders the state the model reports back; it never tracks generation state on its own. That discipline lives in one reducer inapp/lib/state.ts, which projects the
model’s typed state messages into a small SanaState. Because it is fed only by the
useSanaStreamingState hook, the reducer never has to filter by message type, and it reads the
snapshot’s flat, typed fields directly:
app/lib/state.ts
Workspace shell in SanaStreamingApp.tsx subscribes with three typed message hooks:
useSanaStreamingState feeds the reducer, and useSanaStreamingCommandError and
useSanaStreamingGenerationReset handle the two messages that need side effects, imperatively:
app/SanaStreamingApp.tsx
start with no source, resume while not paused); the shell
turns each command_error into a banner that dismisses itself after six seconds.
Every control in the app gates off the reduced SanaState: the file-mode Start button on
state.hasVideo, the mode toggle and clip picker on state.started (the source is fixed once a run
begins), the pause/resume and reset controls on state.started and state.paused. Informational
messages (video_accepted, prompt_accepted, chunk_complete) are not state inputs; whatever they
report also arrives in the next state snapshot, which is the
canonical payload.
The same discipline shapes the commands going out. A command only takes effect once the model echoes
it back in a state snapshot, so the start path doesn’t confirm anything itself: every start flow,
live or file, fires the same two methods and lets the reducer report when generation is running:
app/lib/state.ts
Live mode: editing your webcam
Live mode is the headline feature and the app’s default. Send your webcam to SANA-Streaming by publishing your camera to the model’scamera input track, then setMode({ mode: "live" }) and
start(). Edited frames come back on the main_video track about a second later.
LiveInput.tsx owns the camera acquisition rather than reaching for a declarative webcam component:
app/components/LiveInput.tsx
publish and unpublish from useSanaStreaming() in an effect keyed on the track
and the connection status: it publishes once the session is ready, re-publishes after a reconnect,
and unpublishes on unmount. The typed SDK also ships a declarative <SanaStreamingCameraView> that
acquires and publishes the webcam for you, but it gives no hook to set contentHint, so LiveInput
owns the track and publishes it by hand.
The Start live button calls startGeneration({ setMode, start }, "live") and is disabled until
status === "ready", the publish has resolved, and no generation is running. Switching the mode
toggle to File unmounts LiveInput, which unpublishes the track and stops the webcam, so a mode
switch can’t leave the camera running.
File mode: editing an uploaded clip
File mode trades the camera for an uploaded clip of at least 33 frames. The flow inFileInput.tsx is uploadFile → setVideo → start, with the model’s state as the gate in the
middle:
app/components/FileInput.tsx
video_accepted plus a state snapshot whose has_video is true. The Start edit button is
disabled on !state.hasVideo, which flips when the model accepts the clip, not when the upload
promise resolves. See set_video for the
command contract and File Uploads for what the SDK does with the bytes.
One quirk is worth handling in any client you build: the model sometimes rejects a perfectly valid
clip with a decode failed error. It’s a timing glitch on the model side, not a problem with your
file, and re-sending the same setVideo almost always clears it. So FileInput watches for that
one error with useSanaStreamingCommandError and retries up to twice with the already-uploaded clip
before treating it as real:
app/components/FileInput.tsx
FileInput retries, the
error banner stays silent, so the user never sees a flash of failure for something the app is about
to fix on its own. (The banner skips these by checking an isTransientDecodeFailure helper in
app/lib/state.ts, the same condition FileInput matches above.)
Two behaviors that follow from the model latching its source at start:
- Clip picks are disabled while a run is in progress. A mid-run
setVideowould not take effect until the nextstart, and the UI would show a clip the model isn’t using. Reset first, then pick a new clip. - A file-mode run ends on its own. Once every source frame is transformed, the model emits
generation_completeand returns to idle with the clip, prompt, and seed still staged.startreplays the clip from the top;resetis only needed to swap clips. On completion themain_videotrack freezes on the last transformed frame rather than going dark; the next section covers what the stage does with frozen frames.
The stage
Stage.tsx renders the model output with the typed <SanaStreamingMainVideoView>:
app/components/Stage.tsx
<ReactorView> with track="main_video" pre-bound, and manages the <video> element,
srcObject binding, and browser autoplay quirks for you. Apply your styling to the container around
it, not to the video element it renders.
In file mode with a source loaded, the stage splits into two panes: the original clip on the left,
the transformed stream on the right. The local clip is driven off the reducer state (play when
running, pause when paused, rewind when the source clears) as an approximate sync by design,
with no seeking or drift correction. A status row along the bottom reads running / paused,
currentChunk, and currentPrompt straight off the reduced state.
After a reset the model emits nothing new, so the view would freeze on the last transformed frame;
the shell blacks the stage out until state.running flips back to true. A completed file-mode run
freezes the view the same way, but there the example leaves the last frame visible (the status row
drops back to idle) until the next start or reset.
Steering the prompt mid-stream
Prompts are editing instructions, not scene descriptions: “apply a Van Gogh oil painting style,” not “a Van Gogh painting of a room.”Prompt.tsx is one textarea, one Apply button, and a row of preset
chips, and every path funnels into the same call:
app/components/Prompt.tsx
setPrompt works before start and at any point mid-stream; the model applies it at the next chunk
boundary. A prompt is optional: start without one and the model streams a near-reconstruction until
you set a prompt to steer it. The textarea’s placeholder (“Describe the edit. Changes apply live,
about one chunk later.”) spells out the latency. The active-prompt readout under the button renders
state.currentPrompt, so it reflects what the model is using rather than what was last typed.
The preset chips come from app/lib/examples.ts. They are deliberately short style tags (“Van Gogh
oil painting, swirling brushstrokes, vivid colors”) that show how a whole-frame restyle reads against
a live feed, leaning on the model’s default to carry everything else through. For surgical edits, a
garment swap or a removal where you must spell out what stays fixed, write the fuller instructions
the prompt guide lays out, with its anatomy and a
recipe per edit type.
Playback, seed, and reset
The Input panel (ModeInput.tsx) is phase-aware, driven by the model’s started flag. Before a run
it shows the mode toggle, the active input (webcam or file picker), and the seed field; once
started flips true it keeps the input slot mounted (so live mode keeps publishing) and swaps the
setup controls for Playback.tsx: pause, resume, and reset. Each is a typed method off
useSanaStreaming(), gated on the reduced state:
app/components/Playback.tsx
SeedField.tsx, a setup-phase control that calls setSeed({ seed }) on blur. The
model reads the seed at start, so the same source, prompt, and seed reproduce a run; the field is
keyed to the model-reported seed, so a reset or an external setSeed refreshes it with the model’s
value.
reset does the most work: it aborts the run and clears the model’s source, prompt, and progress,
emitting generation_reset. The shell’s handler (from
The model is the source of truth) mirrors that on the client,
dropping the side-by-side source URL, blacking out the stage, and clearing the prompt draft and file
selection so the UI matches the model.
What’s intentionally left out
The demo covers the connect + edit + steer + capture loop. Clip capture is a shared base-SDK feature, so Recordings covers it, including continuous recording, programmatic capture, and retention. A few other patterns are out of scope, and each is a small addition:
For the full design rationale and the patterns to follow when extending the app, including the
typed-SDK surface and the manual camera publish, read
skill/SKILL.md in
the example repo.