Preserve the state representation
The 14-value state consists of base-relative EE pose, two gripper joint positions, and world-frame base pose. Convert quaternions fromxyzw to axis-angle; never put four quaternion values into a
three-value rotation slot. Preserve the raw simulator coordinate frames and units. The model
performs checkpoint normalization remotely.
Keep a seven-observation client buffer for the reference history settings, sampling t-6, t-4,
t-2, and t. Pair that state history with the server’s actual video samples. See
simulation for the missing transport adapter and receipt
gate.
Execute the simulator action
Once an aligned observation has produced a validated chunk, the upstream execution boundary is:step. This helper only validates
and maps one returned row; it is not a complete rollout loop. The controller owns scaling of EE,
base, torso, and gripper commands. These outputs are decoded controller commands; do not apply
additional checkpoint denormalization or reinterpret them as joint positions.
Execute at most the available 16 rows before replanning. The reference evaluator uses all 16; a
shorter window changes evaluation behavior and should be measured. Update measured state history
before sending a strictly increasing executed-action count. The returned prediction step is a
separate counter and must not be treated as the number of executed simulator steps.