# Python SDK reference

The Python SDK lets your code open sessions on our models, send your robot's data, and
read what you spent. It requires Python 3.8 or newer. Use a virtual environment as
shown in [Quickstart](https://yellowandblack.dev/docs/quickstart.md), then install it with:

```sh
python -m pip install "https://yellowandblack.dev/sdk/robot_inference_client-0.5.3-py3-none-any.whl"
```

You import it as `robot_inference_client`. It has two main classes:

- `Client` connects to the platform. You use it to list models, open sessions, and check
  your usage.
- `PolicySession` is one open session. You use it to send observations and get the next
  moves.

## Client

`example_observation(model="pi05-libero")` is also exported by the SDK. It returns
`(instruction, observation)` from the bundled recorded LIBERO frame, with independent
arrays on each call. It supports `pi05-libero` and `transport-demo`; no download or source
checkout is needed. This replay input has no live capture timestamp and does not qualify
task success. See [Quickstart](https://yellowandblack.dev/docs/quickstart.md) for a complete call.

```python
Client(api_key=None, base_url=None, project_id=None)
```

With no arguments, `Client` connects to the hosted service with the key saved by
`yb setup`. `YB_API_KEY` overrides that key.

It checks the key when you create it. A wrong key raises `AuthorizationError` with the
code `unauthorized`.

Use it in a `with` block, or call `close()` when you're done. Closing a `Client` doesn't
close your sessions: close each one.

| Method | What it does |
|---|---|
| `models(ready=False)` | With `ready=True`, offered models: `model_id`, `status`, `mode` (`warm` or `on_demand`), `rate`, session slots and `contract`. On-demand entries may be idle until a connection starts. Reading this reserves nothing. Without it, every catalog model is returned. |
| `session(model, instruction, max_spend_usd="0", label="Python client", max_action_age_ms=2000, mode="client", idempotency_key=None, transport="sync", connect_timeout=None, on_status=None)` | Waits for model readiness, opens the connection and returns a `PolicySession`. See the options below. |
| `sessions()` | Your project's sessions: state, completed requests, cost, and why each one ended. |
| `model_session(session_id)` | One session, with the model server's live counts while it's open. |
| `close_session(session_id)` | Closes a session by its ID. Safe to call twice. |
| `usage()` | Your credit balance, the credit reserved by open sessions, and your ledger. |
| `regions(model)` and `select_region(model, region="auto")` | For operators only: measure the network delay to each region. |
| `deploy(...)`, `deployment(deployment_id)` and `deployments()` | For operators and developers only: a private model server that one project owns. You don't need these to use the running models. |
| `close()` | Closes the client's connection to the platform. |

### Session options

- `model`: a `model_id` from `models(ready=True)`.
- `connect_timeout`: maximum seconds waiting for on-demand model startup; defaults
  to the model's advertised startup limit. This is separate from individual HTTP
  and worker handshake timeouts.
  Temporary `database_busy` or `executor_busy` status responses are polled again
  within this same limit. Admission and inference requests are not replayed.
  A readiness response received at or after the startup limit is treated as
  `model_start_timeout`; the SDK requests cancellation without renewing a grant.
- `on_status`: optional callback receiving `"starting"` and `"ready"`, for example
  `on_status=print`. Loading and waiting are included; they are not inference charges.
  Ctrl-C or a startup timeout requests cancellation. A startup `PlatformError`
  includes `session_id`, `cleanup_requested` and `cleanup_confirmed` in its context;
  if requested cleanup was not confirmed, call `close_session` or check its status.
  Server-side startup and attach deadlines still expire the admission.
- `instruction`: the task in words, such as `"put the bowl on the plate"`. Every
  observation's `prompt` must match it. Change it with `reset`.
- `max_spend_usd`: the most this session may cost, as a string or `Decimal`, such as
  `"5"`. A float is refused. It's reserved from your credit while the session is open,
  and it also caps the number of requests. `"0"` works only on the free sandbox.
- `max_action_age_ms`: the deadline, meaning how old an answer may be when it reaches you,
  counted from when the cameras took the pictures. The default is 2000, and the range is
  20 to 30000. Older answers are dropped. For example, `max_action_age_ms=500`
  drops any answer older than half a second. An answer the model server drops as late
  is free. An answer the server sent in time, but that reached you late, is charged.
- `idempotency_key`: if your code retries after a crash, pass the same value each time,
  so a retry never opens a second session. By default, each call gets a new value. While
  that session is open, the same value returns it. Once it has ended, you get an error
  such as `session_closed`: open a new session with a new value.
  If the admission response is lost, the `PlatformError` context includes the generated
  `idempotency_key`; use it to recover that admission. A failed connection attempt on
  an already active replay does not automatically close the other connection.
- `transport`: `"sync"` (the default) or `"owned"`. With `"owned"`, every network wait
  gives up at the request's deadline, so a stuck network can't freeze your robot's loop.
  Use `"owned"` for a control loop on a real robot. It needs `websockets` 13 or newer.
  The default is fine for simulators.
- `image_codec`: `"jpeg"` (the default) sends camera images as JPEG, quality 90, which is
  14 times fewer bytes than raw pixels. π0.5 succeeded on 30 of 30 LIBERO tasks with it,
  against 29 of 30 with raw pixels (measured). `"raw"` sends exact pixels, for example to
  compare against a local model bit for bit. A model server that doesn't accept JPEG gets
  raw pixels automatically. `jpeg_quality` (50 to 100, default 90) sets the quality.
- `udp`: `"auto"` (the default), `"on"` or `"off"`. With `"auto"`, observations go over
  UDP when the model server offers it and a test packet gets through as the session
  opens; otherwise they go over the WebSocket. UDP doesn't wait for lost packets to be
  sent again: spare packets rebuild most losses on arrival, and the server asks for the
  rest. On a link with a 20 ms round trip that lost 1% of packets each way, the slowest 1%
  of requests took 119 ms over UDP and 356 ms over the WebSocket (measured, including
  86 ms of model time). With no loss the two tie. Answers come back both ways, and you
  get whichever arrives first. If two UDP requests in a row get no answer at all, the
  session switches to the WebSocket by itself, and tries UDP again every 30 seconds.
  `"on"` raises `fastpath_unavailable` when UDP can't be used as the session opens;
  `"off"` never tries UDP. The `YB_UDP` variable sets the default. UDP needs outbound UDP to the model server's port; see
  [Security](https://yellowandblack.dev/docs/security.md) for how it's encrypted.

One robot's open session on one model. Use one per robot, from one thread at a time: its
methods must not run at the same time.

| Member | What it does |
|---|---|
| `infer(observation, captured_ns=None)` | Sends one observation (camera images, arm state and task) and waits for the answer. Returns a dict: `actions` (the next moves, one row per step, as a NumPy array), `client_action_age_ms` (how old the answer was on arrival), `sequence` and `epoch` (which request and which episode it answers), `model_revision` (which model version answered) and `timing`. Raises `InferenceError` if the request fails. |
| `submit(observation, captured_ns=None)` | Sends one observation without waiting, and returns its `sequence` number. Collect the answer with `result`. Send the next observation while the robot still executes the current actions, so the model's time is hidden behind the motion. The model server keeps only the newest waiting observation per session, so submitting again before the last one started replaces it (`superseded`). |
| `result(sequence, timeout=None)` | Waits for the answer to a submitted observation, until its deadline, and returns the same dict as `infer`. With `timeout` (seconds) it waits less and raises `TimeoutError`, leaving the request pending, so your loop can do other work and ask again. |
| `capture_time_ns()` | The SDK's clock, in nanoseconds. Read it when the cameras take the pictures, and pass it as `captured_ns`. |
| `network()` | How the session reaches the model server: `path` (`"udp"` or `"websocket"`), `udp_round_trip_ms` (measured as the session opened), `note` (why UDP isn't in use, if it isn't), `answers` (how many answers arrived first by each path) and `udp` (packet counts, including resends). |
| `reset(instruction=None)` | Starts a new episode (a new attempt at the task), optionally with a new task. Answers meant for the old episode are dropped. |
| `status()` | The platform's view of the session: state, completed requests, cost so far, and the model server's counts. |
| `renew_credentials()` | Gets a fresh session pass now. You rarely need it: the SDK renews the pass once half its life has passed, and through a short platform outage it keeps trying until the pass is about to expire. |
| `report()` | The model server's report on the session: counts and timings. |
| `flush_telemetry()` | Sends up to 100 recent timings measured on your side, to help find problems. It never changes what you pay. |
| `close()` | Ends the session and settles its cost. Safe to call twice. Leaving a `with` block calls it. |
| `closure` | Set by `close()`: `completed` (requests answered), `charged_usd`, `settlement_source` (see [Billing](https://yellowandblack.dev/docs/billing.md)), `close_confirmed` and `close_reason`. `close_reason` is `customer_closed` when your code closed the session, or `worker:` plus the model server's reason when it had already ended the session, such as `worker:too_many_expired`. It's `None` if the platform couldn't be reached; the session then ends on its own, and nothing extra is charged. If the model server couldn't be reached, `closure` shows the state `closing` and `charged_usd` 0 until the platform settles it. |
| `assignment` | The platform's record of the session: model, version, the model server it runs on (`generation`), price, spending cap and request limit. |
| `id`, `model_id`, `model_revision`, `generation` | Which session this is, which model version it runs, and which model server it's on. |

## Errors

Every error has a `code`, such as `capacity`. [Errors](https://yellowandblack.dev/docs/errors.md) says what each code
means and how to fix it.

| Class | Raised when | Extra members |
|---|---|---|
| `PlatformError` | a call to the platform fails. Codes not listed below, such as `terms_required` or `email_unverified`, raise a plain `PlatformError`. | `status` (the HTTP status), `context` (extra details), `retryable` |
| `CapacityError` | `no_capacity`, `capacity`, `session_limit`, `account_session_limit` or `rate_limit` (a kind of `PlatformError`) | the same |
| `AuthorizationError` | `unauthorized`, `forbidden`, `not_found` or `rate_changed` | the same |
| `SpendError` | `insufficient_credit`, `spend_too_small` or `spend_limit` | the same |
| `InferenceError` | `infer` or `reset` fails: for example a late answer (`expired`), a closed session, or a dropped connection (`connection_lost`). | none |

Every error also has `code`, `message`, `fix`, `retry` and `docs_url`. `retry` says what
to do next: `retry`, `fix`, `next_request`, `new_session` or `none`. The table on
[Errors](https://yellowandblack.dev/docs/errors.md) explains each. Printing an error shows all of these.

For example, to retry when the model is full, waiting longer each time:

```python
import time

from robot_inference_client import CapacityError


def open_session(client, instruction, attempts=5):
    for attempt in range(attempts):
        try:
            return client.session("pi05-libero", instruction, max_spend_usd="5")
        except CapacityError as error:
            if not error.retryable or attempt == attempts - 1:
                raise
            time.sleep(2**attempt)  # nothing is charged while you wait
```

## Environment variables

| Variable | Meaning |
|---|---|
| `YB_URL` | Another copy of the platform to use instead of the hosted service. You don't need it otherwise. |
| `YB_API_KEY` | Your key, instead of the one saved by `yb setup`. For servers and CI. |
| `YB_PROJECT_ID` | Which project to use, when one key can see more than one. |
| `YB_ALLOW_INSECURE_HTTP` | Set to `1` to allow plain `http` to another machine on a private network you control. Without it, the SDK sends keys only over `https`, or to this machine. |
| `YB_UDP` | The default for the `udp` session option: `auto`, `on` or `off`. |

## Robot safety

The SDK checks that each answer belongs to the request you sent, and that it's fresh when
it arrives. It can't know whether the moves still fit the world when your robot makes
them. Your robot's controller must check that sensor data is fresh, that each move still
makes sense, and when to stop.
