reactor/dreamzero-yam-molmoact2 through the SDK and keep the session open while
exchanging observations and predictions. A camera track is a named stream of successive images
from one camera; publishing it attaches that stream to the session.
API at a glance
Chunks stream as observations arrive; there is no separate predict command or execution
acknowledgement. Your application validates the output and executes it through its own controller.
New to Reactor? How the API works explains sessions, tracks,
commands, and the difference between an SDK message and the data returned by the example helper. For
a complete script, use Get your first actions.
Camera inputs
Publishtop, left, and right as RGB uint8 arrays shaped (H, W, 3) after READY. They map
respectively to top_camera-images-rgb, left_camera-images-rgb, and right_camera-images-rgb in
the checkpoint. The evaluation transform uses (176, 320) images; the endpoint resizes incoming
views. Preserve view identity, mounting, and field of view.
Commands
The zero joint lists above illustrate payload shape only. Real control requires measured state for
both arms in the checkpoint’s native joint order; these are not Cartesian poses.
Action output
Theaction_chunk message’s data contains:
Each row is laid out as follows:
Joint targets are absolute radians for each arm whose measured state is supplied. An omitted arm
state defaults to zero and leaves that arm’s output relative to the zero anchor. Every row is
referenced to the chunk’s anchor state, not the preceding row. Do not cumulatively sum rows or add
joint state a second time. Gripper columns are continuous normalized closure; validate and map them
locally. Action width alone does not verify a robot’s joint order or calibration.
The API supplies no per-row execution timestamp or controller rate. Establish the cadence from your
YAM data and controller integration; the chunk arrival rate is not the execution rate.
Capture-time alignment
Capture-time pairing ships disabled. Sendset_pair_by_capture_time explicitly to enable it;
pair_by_capture_time_accepted reports the selected value. The override takes effect from the next
chunk and persists across reset. Unset session state follows the deployment default.
When enabled and stamps are available, selection trims camera histories near the slowest view’s
newest capture timestamp, with approximately 33 ms tolerance, before selecting each window. It can
wait for usable history. If a newest frame lacks a timestamp, selection falls back to arrival order.
This changes the observations seen by the model, so evaluate it on your capture setup.
obs_capture_time_us describes the sender’s clock domain. view_skew_us requires stamps on all
three newest consumed frames; obs_capture_time_us can still be present when skew is null. Neither
field certifies hardware synchronization. Only compute observation age against a compatible clock;
subtracting an unrelated server wall clock does not measure end-to-end latency.
predicted_video carries a composite preview when enabled and black frames otherwise. It is
optional diagnostic output and adds decoding work.
Streaming and episode state
The first chunk uses one frame per camera. Subsequent chunks use a rolling window of up to four frames per view, padding short histories with the oldest frame. Inference waits until every camera has delivered a new frame since the preceding chunk. Missing one view can stall the whole loop.obs_seq is the largest server ingest counter among frames selected for a chunk. It is not a client
request ID, capture timestamp, or proof that a chunk includes a particular just-published frame. A
chunk already in flight can arrive after you update the observations. Rejecting counters older than
those already consumed prevents replay, but does not establish exact frame pairing.
set_prompt returns prompt_accepted. A different prompt re-anchors the causal cache; it is not an
episode reset. reset returns episode_reset, clears the prompt and measured robot state, and
restarts the episode counters when inference restarts. The first action has chunk_index: 0; its
obs_seq need not be zero because frames have already been ingested. After reset, discard pending
actions and send fresh state, frames, and a prompt. A model reset does not stop the robot.