Skip to main content
Connect to reactor/dreamzero-yam-molmoact2 through the SDK and keep the session open while exchanging observations and predictions. A camera track is a named stream of successive images from one camera; publishing it attaches that stream to the session.

API at a glance

Chunks stream as observations arrive; there is no separate predict command or execution acknowledgement. Your application validates the output and executes it through its own controller. New to Reactor? How the API works explains sessions, tracks, commands, and the difference between an SDK message and the data returned by the example helper. For a complete script, use Get your first actions.

Camera inputs

Publish top, left, and right as RGB uint8 arrays shaped (H, W, 3) after READY. They map respectively to top_camera-images-rgb, left_camera-images-rgb, and right_camera-images-rgb in the checkpoint. The evaluation transform uses (176, 320) images; the endpoint resizes incoming views. Preserve view identity, mounting, and field of view.

Commands

The zero joint lists above illustrate payload shape only. Real control requires measured state for both arms in the checkpoint’s native joint order; these are not Cartesian poses.

Action output

The action_chunk message’s data contains: Each row is laid out as follows: Joint targets are absolute radians for each arm whose measured state is supplied. An omitted arm state defaults to zero and leaves that arm’s output relative to the zero anchor. Every row is referenced to the chunk’s anchor state, not the preceding row. Do not cumulatively sum rows or add joint state a second time. Gripper columns are continuous normalized closure; validate and map them locally. Action width alone does not verify a robot’s joint order or calibration. The API supplies no per-row execution timestamp or controller rate. Establish the cadence from your YAM data and controller integration; the chunk arrival rate is not the execution rate.

Capture-time alignment

Capture-time pairing ships disabled. Send set_pair_by_capture_time explicitly to enable it; pair_by_capture_time_accepted reports the selected value. The override takes effect from the next chunk and persists across reset. Unset session state follows the deployment default. When enabled and stamps are available, selection trims camera histories near the slowest view’s newest capture timestamp, with approximately 33 ms tolerance, before selecting each window. It can wait for usable history. If a newest frame lacks a timestamp, selection falls back to arrival order. This changes the observations seen by the model, so evaluate it on your capture setup. obs_capture_time_us describes the sender’s clock domain. view_skew_us requires stamps on all three newest consumed frames; obs_capture_time_us can still be present when skew is null. Neither field certifies hardware synchronization. Only compute observation age against a compatible clock; subtracting an unrelated server wall clock does not measure end-to-end latency. predicted_video carries a composite preview when enabled and black frames otherwise. It is optional diagnostic output and adds decoding work.

Streaming and episode state

The first chunk uses one frame per camera. Subsequent chunks use a rolling window of up to four frames per view, padding short histories with the oldest frame. Inference waits until every camera has delivered a new frame since the preceding chunk. Missing one view can stall the whole loop. obs_seq is the largest server ingest counter among frames selected for a chunk. It is not a client request ID, capture timestamp, or proof that a chunk includes a particular just-published frame. A chunk already in flight can arrive after you update the observations. Rejecting counters older than those already consumed prevents replay, but does not establish exact frame pairing. set_prompt returns prompt_accepted. A different prompt re-anchors the causal cache; it is not an episode reset. reset returns episode_reset, clears the prompt and measured robot state, and restarts the episode counters when inference restarts. The first action has chunk_index: 0; its obs_seq need not be zero because frames have already been ingested. After reset, discard pending actions and send fresh state, frames, and a prompt. A model reset does not stop the robot.