reactor/dreamzero through the SDK and keep the session open while exchanging
observations and predictions. A camera track is a named stream of successive images from one
camera; publishing it attaches that stream to the session.
API at a glance
Chunks stream as observations arrive; there is no separate predict command or execution
acknowledgement. Your application validates the output and executes it through its own controller.
New to Reactor? How the API works explains sessions, tracks,
commands, and the difference between an SDK message and the data returned by the example helper. For
a complete script, use Get your first actions.
Camera inputs
Publish RGBuint8 arrays shaped (H, W, 3) on all three named tracks after READY.
The model resizes input images to
(180, 320). RoboLab’s default second exterior image is black;
keep that mapping explicit. Reversing the exterior tracks can leave the primary view black without
producing a protocol error.
Commands
Callsession.send(command, payload) with these payloads:
Joint order is the DROID arm order, corresponding to
panda_joint1 through panda_joint7 in
RoboLab. Send finite measured values, not the preceding commanded targets. State and camera tracks
travel asynchronously; sending state first is useful ordering, not an atomic observation
transaction.
Action output
Theaction_chunk envelope contains this data payload:
When measured state is supplied, the model’s inverse transform produces absolute joint targets in
radians. Do not add the measured joints again or integrate rows as velocities. With unset joint
state, zeros are substituted and the returned joints instead reflect relative predictions. Gripper
output is continuous normalized closure. Validate its range locally before mapping it to your
controller; the endpoint does not guarantee all predictions lie within physical limits.
The reference simulator uses 15 Hz control: 24 rows represent 1.6 seconds of targets if all are
executed. That control rate is independent of inference chunk arrival rate.
predicted_video is a composite video output. It carries black frames while preview is disabled;
enabling set_video_preview adds decoding work. Preview frames are predictions, not sensor input.
Streaming and episode state
The first chunk uses one frame per camera. Subsequent chunks use a rolling window of up to four frames per view, padding short histories with the oldest frame. Inference waits until every camera has delivered a new frame since the preceding chunk. Missing one view can stall the whole loop.obs_seq is the largest server ingest counter among frames selected for a chunk. It is not a client
request ID, capture timestamp, or proof that a chunk includes a particular just-published frame. A
chunk already in flight can arrive after you update the observations. Rejecting counters older than
those already consumed prevents replay, but does not establish exact frame pairing.
set_prompt returns prompt_accepted. A different prompt re-anchors the causal cache; it is not an
episode reset. reset returns episode_reset, clears the prompt and measured robot state, and
restarts the episode counters when inference restarts. The first action has chunk_index: 0; its
obs_seq need not be zero because frames have already been ingested. After reset, discard pending
actions and send fresh state, frames, and a prompt. A model reset does not stop the robot.