Skip to main content
Connect to reactor/cosmos-nano-policy-droid through the SDK and keep the session open while exchanging observations and predictions. A camera track is a named stream of successive images from one camera; publishing it attaches that stream to the session.

API at a glance

The first chunk arrives once valid observations and a task are available. After executing it, send updated observations and report which prediction you executed to receive the next chunk. Your application validates the output and executes it through its own controller. New to Reactor? How the API works explains sessions, tracks, commands, and the difference between an SDK message and the data returned by the example helper. For a complete script, use Get your first actions.

Camera tracks

Publish all three tracks after READY. Supply RGB uint8 arrays shaped (H, W, 3). These are the reference simulator client’s dimensions, not schema-enforced resolutions. The first-action fixture uses (180, 320, 3) for all views. The model composes and resizes the images internally; preserve camera identity and field of view when adapting another rig. The example tracks publish at 15 fps. The endpoint retains the newest frame from each view. Video delivery and state commands are asynchronous; replies contain no frame timestamp or observation identifier. A settling delay reduces stale-frame risk but does not guarantee synchronized observations.

Commands

proprio_json contains row lists, even for a single observation:
Send measured positions, not the last requested targets. Joint columns follow DROID’s seven arm joints (panda_joint1 through panda_joint7 in RoboLab), in radians. Gripper state is normalized closure: 0 open, 1 closed. For history, use matching, nonempty (N, 7) and (N, 1) lists with the current sample last. Validate finite values locally; the endpoint does not enforce every physical constraint.

Action reply

The action_prediction message’s data object contains: Each row is [q1, q2, q3, q4, q5, q6, q7, gripper]. Joint values are absolute position targets in radians, not velocities or increments. Gripper values use the same closure convention as the input. The raw prediction is continuous and is not guaranteed to stay within [0, 1]. RoboLab thresholds gripper output at > 0.5 to close; <= 0.5 opens. The endpoint does not apply this threshold for you. The reference simulation executes rows at 15 Hz: one chunk represents about 2.13 seconds of motion. This is a control cadence, not a promise of 15 inference replies per second or an inference latency bound. No generated video is returned.

Advance the loop

  1. Send the task, publish all camera views, then send valid proprioception. The first prediction requires no acknowledgement and returns step: 0.
  2. Execute the chunk in your application. Publish the next observations and send updated proprioception.
  3. Send set_executed_step_json with a JSON string containing the received step and the rows actually executed: {"step": 0, "action": [[...], ...]}. The next reply has step: 1.
  4. Keep one chunk outstanding. Validate shape, finite values, and the expected next counter.
The server opens the gate when the acknowledged counter exceeds the last acknowledged counter; it does not verify execution or compare the echoed rows with its prediction. Repeating the same counter does not request a retransmission. Do not invent a larger counter to recover a timeout.

Task changes and reset

A changed task applies to the next prediction; it does not bypass the acknowledgement gate. reset restarts the counter and clears retained camera frames, but does not clear the input state fields. Old proprioception and acknowledgements can therefore still be present. Use a fresh session after an uncertain timeout or at an episode boundary when you need unambiguous pairing. A model reset never resets or stops the physical robot.