Skip to main content
Connect to reactor/dreamzero through the SDK and keep the session open while exchanging observations and predictions. A camera track is a named stream of successive images from one camera; publishing it attaches that stream to the session.

API at a glance

Chunks stream as observations arrive; there is no separate predict command or execution acknowledgement. Your application validates the output and executes it through its own controller. New to Reactor? How the API works explains sessions, tracks, commands, and the difference between an SDK message and the data returned by the example helper. For a complete script, use Get your first actions.

Camera inputs

Publish RGB uint8 arrays shaped (H, W, 3) on all three named tracks after READY. The model resizes input images to (180, 320). RoboLab’s default second exterior image is black; keep that mapping explicit. Reversing the exterior tracks can leave the primary view black without producing a protocol error.

Commands

Call session.send(command, payload) with these payloads: Joint order is the DROID arm order, corresponding to panda_joint1 through panda_joint7 in RoboLab. Send finite measured values, not the preceding commanded targets. State and camera tracks travel asynchronously; sending state first is useful ordering, not an atomic observation transaction.

Action output

The action_chunk envelope contains this data payload: When measured state is supplied, the model’s inverse transform produces absolute joint targets in radians. Do not add the measured joints again or integrate rows as velocities. With unset joint state, zeros are substituted and the returned joints instead reflect relative predictions. Gripper output is continuous normalized closure. Validate its range locally before mapping it to your controller; the endpoint does not guarantee all predictions lie within physical limits. The reference simulator uses 15 Hz control: 24 rows represent 1.6 seconds of targets if all are executed. That control rate is independent of inference chunk arrival rate. predicted_video is a composite video output. It carries black frames while preview is disabled; enabling set_video_preview adds decoding work. Preview frames are predictions, not sensor input.

Streaming and episode state

The first chunk uses one frame per camera. Subsequent chunks use a rolling window of up to four frames per view, padding short histories with the oldest frame. Inference waits until every camera has delivered a new frame since the preceding chunk. Missing one view can stall the whole loop. obs_seq is the largest server ingest counter among frames selected for a chunk. It is not a client request ID, capture timestamp, or proof that a chunk includes a particular just-published frame. A chunk already in flight can arrive after you update the observations. Rejecting counters older than those already consumed prevents replay, but does not establish exact frame pairing. set_prompt returns prompt_accepted. A different prompt re-anchors the causal cache; it is not an episode reset. reset returns episode_reset, clears the prompt and measured robot state, and restarts the episode counters when inference restarts. The first action has chunk_index: 0; its obs_seq need not be zero because frames have already been ingested. After reset, discard pending actions and send fresh state, frames, and a prompt. A model reset does not stop the robot.