Skip to main content
LingBot-VA LIBERO takes an external camera, a wrist camera, and a language task. It returns a chunk of end-effector delta commands and uses the executed actions and observed video to predict the next chunk. It has no proprioception input or generated-video output on this endpoint. The control rate describes execution, not inference throughput. Replies arrive after inference and transport delay; the reference simulator pauses between chunks. Follow Get your first actions, then simulation or robot integration. The reference defines the wire and action semantics.