Skip to main content
Connect to reactor/xr1-robocasa365 through the SDK and keep the session open while exchanging observations and predictions. A camera track is a named stream of successive images from one camera; publishing it attaches that stream to the session.

API at a glance

Every prediction, including the first, requires a new execution-progress message and enough complete camera sets. Your application validates the output and executes it through its own controller. New to Reactor? How the API works explains sessions, tracks, commands, and the difference between an SDK message and the data returned by the example helper. For a complete script, use Get your first actions.

Camera tracks

Use RGB uint8 images; the upstream benchmark renders (256, 256, 3). The server applies its center crop and model image preprocessing, so do not apply the checkpoint’s crop a second time. Publish each captured three-view observation as a set. The server pairs the nth arrival on each track; it does not match camera capture timestamps. Dropping or duplicating frames on one track can silently misalign the history.

Commands

JSON-bearing fields are strings. The parser accepts exactly four equal-width rows of 1–60 finite numbers, but checkpoint-compatible inputs use 14 columns. Parser acceptance alone does not establish a valid robot state.

State layout

The upstream evaluator constructs each row as follows. Slice ends are exclusive. The server pads columns 14:60 and applies checkpoint state normalization. Send the simulator’s measurements without training-statistic normalization. Do not collapse the two gripper joint values into a single open/closed action value.
The release 0.2.2 schema’s state description labels the 14 values as left/right arm poses. That prose does not match this single-arm checkpoint’s upstream evaluator. Use the EE/gripper/base layout above.

History and readiness

For the reference history settings, retain seven consecutive observations and select indices t-6, t-4, t-2, and t, oldest first. At startup, clamp missing indices to the first observation. Use the same sampling for state and all camera views. The serving implementation maintains its own camera buffer and samples four frames at interval two by default. It accepts state history already sampled by the client. Do not publish four pre-sampled frames and assume the server will preserve their spacing: that can sample the history twice. Every prediction, including the first, needs a fresh execution echo and enough complete camera sets received. The cumulative minimum is four sets per accepted echo: 4 for the first prediction, 8 for the second, and so on. The echo’s integer may count executed actions (0, 16, 32, …); it is not used as a frame count. A compatible simulator bridge must align this receipt gate with history sampling. See simulation.

Action reply

The message envelope is {"type": "action_prediction", "data": {...}}. Only action[:, :12] is used by the benchmark. Pass each selected row through RoboCasa’s convert_action to map the controller commands to the simulator action dictionary. These are the simulator controller’s scaled commands, not absolute joint targets or raw Cartesian metres/radians. Checkpoint action denormalization has already run. RoboCasa’s conversion and controller mapping own the physical scaling and mode interpretation.

Replanning

Execute the chosen window (upstream uses all 16 rows), collect aligned observations, update state_history_json, and then send an increasing execution echo. Allow one outstanding request. Changing state or task alone does not trigger another prediction. This variant has no prefix conditioning or prefix_rows output. Track execution separately from response step. At an episode boundary, stop execution, discard queued responses, and reset local history. Reconnect for strict isolation: reset clears the server’s history and counters but does not replace all stored input fields or guarantee that old media has left the transport.