# Robot data formats

Each model expects your robot's data in one exact format, and replies in one exact format. This page shows both, for every model: the names to use, the size of each camera image and list of numbers, and the shape of the moves that come back.

- The model server checks your data, and refuses data in the wrong format. With the SDK, one observation in the wrong format ends the session with `invalid_observation`.
- The model server always replies in the format shown here, and the SDK checks each answer.
- `yb models --ready` shows each model's format, under `contract`.

Camera images are 224 × 224 color (RGB) arrays of `uint8`. The robot's state is a list of ordinary numbers (no NaN or infinity). The `prompt` is the task in words, and must match the session's task; change it with `reset`.

## `openpi-libero-v1`

Robot: LIBERO simulator (Franka, robosuite). How it moves: end-effector delta + gripper, about 10 steps per second.

| Send | Type | Shape | Meaning |
|---|---|---|---|
| `observation/image` | uint8 | 224 × 224 × 3 | Main camera, RGB |
| `observation/wrist_image` | uint8 | 224 × 224 × 3 | Wrist camera, RGB |
| `observation/state` | float | 8 | eef_x, eef_y, eef_z, axis_angle_x, axis_angle_y, axis_angle_z, gripper_0, gripper_1 |
| `prompt` | str | | The session's task, in words |

Returns `actions`: an array of numbers of shape (10, 7): one row per step, one column per number. Columns: dx, dy, dz, drot_x, drot_y, drot_z, gripper. Carry out the first few steps, then send a new observation.

Recorded September 25 model test (historical): the π0.5 LIBERO model was tested on a rented A100 GPU on 2026-09-25 (86 ms per request). Repeated runs on the same input give slightly different actions (largest difference about 0.004); we are still working to make them identical.

For later Modal serving and startup measurements, see [Performance and availability](https://yellowandblack.dev/docs/faq.md#how-fast-is-it). These dated records do not qualify a physical robot or guarantee current capacity.

```python
import numpy as np

observation = {
    "observation/image": np.zeros((224, 224, 3), dtype=np.uint8),
    "observation/wrist_image": np.zeros((224, 224, 3), dtype=np.uint8),
    "observation/state": np.zeros(8, dtype=np.float32),
    "prompt": "put the bowl on the plate",
}
```

## `openpi-droid-v1`

Robot: DROID (Franka Panda, joint-velocity control). How it moves: joint velocity + gripper position, about 15 steps per second.

| Send | Type | Shape | Meaning |
|---|---|---|---|
| `observation/exterior_image_1_left` | uint8 | 224 × 224 × 3 | Exterior camera, RGB |
| `observation/wrist_image_left` | uint8 | 224 × 224 × 3 | Wrist camera, RGB |
| `observation/joint_position` | float | 7 | q1, q2, q3, q4, q5, q6, q7 |
| `observation/gripper_position` | float | 1 | gripper_position |
| `prompt` | str | | The session's task, in words |

Returns `actions`: an array of numbers of shape (15, 8): one row per step, one column per number. Columns: dq1, dq2, dq3, dq4, dq5, dq6, dq7, gripper_position. Carry out the first few steps, then send a new observation.

Recorded September 25 model test (historical): the π0.5 DROID model was tested on a rented A100 GPU on 2026-09-25 (87 ms per request). It has not yet been tested on a real DROID robot.

For later Modal serving and startup measurements, see [Performance and availability](https://yellowandblack.dev/docs/faq.md#how-fast-is-it). These dated records do not qualify a physical robot or guarantee current capacity.

```python
import numpy as np

observation = {
    "observation/exterior_image_1_left": np.zeros((224, 224, 3), dtype=np.uint8),
    "observation/wrist_image_left": np.zeros((224, 224, 3), dtype=np.uint8),
    "observation/joint_position": np.zeros(7, dtype=np.float32),
    "observation/gripper_position": np.zeros(1, dtype=np.float32),
    "prompt": "put the bowl on the plate",
}
```
