Robot data formats
Each model expects your robot's data in one exact format, and replies in one exact format. This page shows both, for every model: the names to use, the size of each camera image and list of numbers, and the shape of the moves that come back.
- The model server checks your data, and refuses data in the wrong format. With the SDK, one observation in the wrong format ends the session with
invalid_observation. - The model server always replies in the format shown here, and the SDK checks each answer.
yb models --readyshows each model's format, undercontract.
Camera images are 224 × 224 color (RGB) arrays of uint8. The robot's state is a list of ordinary numbers (no NaN or infinity). The prompt is the task in words, and must match the session's task; change it with reset.
openpi-libero-v1
Robot: LIBERO simulator (Franka, robosuite). How it moves: end-effector delta + gripper, about 10 steps per second.
| Send | Type | Shape | Meaning |
|---|---|---|---|
observation/image |
uint8 | 224 × 224 × 3 | Main camera, RGB |
observation/wrist_image |
uint8 | 224 × 224 × 3 | Wrist camera, RGB |
observation/state |
float | 8 | eef_x, eef_y, eef_z, axis_angle_x, axis_angle_y, axis_angle_z, gripper_0, gripper_1 |
prompt |
str | The session's task, in words |
Returns actions: an array of numbers of shape (10, 7): one row per step, one column per number. Columns: dx, dy, dz, drot_x, drot_y, drot_z, gripper. Carry out the first few steps, then send a new observation.
Recorded September 25 model test (historical): the π0.5 LIBERO model was tested on a rented A100 GPU on 2026-09-25 (86 ms per request). Repeated runs on the same input give slightly different actions (largest difference about 0.004); we are still working to make them identical.
For later Modal serving and startup measurements, see Performance and availability. These dated records do not qualify a physical robot or guarantee current capacity.
import numpy as np
observation = {
"observation/image": np.zeros((224, 224, 3), dtype=np.uint8),
"observation/wrist_image": np.zeros((224, 224, 3), dtype=np.uint8),
"observation/state": np.zeros(8, dtype=np.float32),
"prompt": "put the bowl on the plate",
}
openpi-droid-v1
Robot: DROID (Franka Panda, joint-velocity control). How it moves: joint velocity + gripper position, about 15 steps per second.
| Send | Type | Shape | Meaning |
|---|---|---|---|
observation/exterior_image_1_left |
uint8 | 224 × 224 × 3 | Exterior camera, RGB |
observation/wrist_image_left |
uint8 | 224 × 224 × 3 | Wrist camera, RGB |
observation/joint_position |
float | 7 | q1, q2, q3, q4, q5, q6, q7 |
observation/gripper_position |
float | 1 | gripper_position |
prompt |
str | The session's task, in words |
Returns actions: an array of numbers of shape (15, 8): one row per step, one column per number. Columns: dq1, dq2, dq3, dq4, dq5, dq6, dq7, gripper_position. Carry out the first few steps, then send a new observation.
Recorded September 25 model test (historical): the π0.5 DROID model was tested on a rented A100 GPU on 2026-09-25 (87 ms per request). It has not yet been tested on a real DROID robot.
For later Modal serving and startup measurements, see Performance and availability. These dated records do not qualify a physical robot or guarantee current capacity.
import numpy as np
observation = {
"observation/exterior_image_1_left": np.zeros((224, 224, 3), dtype=np.uint8),
"observation/wrist_image_left": np.zeros((224, 224, 3), dtype=np.uint8),
"observation/joint_position": np.zeros(7, dtype=np.float32),
"observation/gripper_position": np.zeros(1, dtype=np.float32),
"prompt": "put the bowl on the plate",
}