Skip to main content
This checkpoint targets the RoboCasa365 PandaOmron simulator. A physical Franka arm or mobile base does not become compatible because it has similar actuators. No validated physical robot driver is provided here. Start with a simulator adapter and the checkpoint’s observation/controller conventions.

Preserve the state representation

The 14-value state consists of base-relative EE pose, two gripper joint positions, and world-frame base pose. Convert quaternions from xyzw to axis-angle; never put four quaternion values into a three-value rotation slot. Preserve the raw simulator coordinate frames and units. The model performs checkpoint normalization remotely. Keep a seven-observation client buffer for the reference history settings, sampling t-6, t-4, t-2, and t. Pair that state history with the server’s actual video samples. See simulation for the missing transport adapter and receipt gate.

Execute the simulator action

Once an aligned observation has produced a validated chunk, the upstream execution boundary is:
Pass this dictionary to your configured RoboCasa environment’s step. This helper only validates and maps one returned row; it is not a complete rollout loop. The controller owns scaling of EE, base, torso, and gripper commands. These outputs are decoded controller commands; do not apply additional checkpoint denormalization or reinterpret them as joint positions. Execute at most the available 16 rows before replanning. The reference evaluator uses all 16; a shorter window changes evaluation behavior and should be measured. Update measured state history before sending a strictly increasing executed-action count. The returned prediction step is a separate counter and must not be treated as the number of executed simulator steps.

Control ownership

Keep the action executor local and bounded: validate outputs, enforce controller limits, and define safe behavior for stale observations, missed deadlines, disconnects, and simulator termination. Discard unused actions when replanning or changing episodes. Reset both sides’ history; reconnect when old media must not cross the episode boundary. Moving to physical hardware additionally requires an appropriate checkpoint, calibrated cameras, kinematic/frame mapping, controller scaling, execution timing, and a robot safety system. The synthetic quickstart proves none of those conditions.

Observation timing

Camera tracks and state commands arrive independently. Timestamp observations at capture and set limits for their age and cross-camera skew; a reply counter does not establish capture alignment. Publish actual observations rather than treating repeated images as new measurements. Measure capture-to-execution age and controller queue depth. Validate the complete loop in a simulator or replay harness before adapting it to hardware.