# Yellow and Black > Yellow and Black runs AI models for robots on its GPUs, starting with vision-language-action (VLA) models such as π0.5. Open a session and run inference for a real robot or a simulation. --- # Quickstart Yellow and Black runs AI models for robots on our GPUs. Today it serves vision-language-action (VLA) models, such as π0.5. It lets you: - Open a session; we prepare the selected model and connect you when it is ready - Run inference for a real robot or a simulation, with one of our models ## Set up 1. Create an account at [https://yellowandblack.dev/platform](https://yellowandblack.dev/platform), accept the service terms, and verify your email to receive the one-time **$5 trial credit**. Then create an API key in the **API keys** tab. 2. Use a virtual environment for your project. If you already have one, activate it. Otherwise, create one in your project folder. On macOS or Linux: ```sh python3 -m venv .venv source .venv/bin/activate ``` On Windows PowerShell: ```powershell py -m venv .venv .\.venv\Scripts\Activate.ps1 ``` If PowerShell blocks activation, use `.\.venv\Scripts\python.exe` instead of `python` and `.\.venv\Scripts\yb.exe` instead of `yb` in the commands below. Activate the environment again when you open a new terminal. Install the SDK and add your key. Paste the key when asked, and answer `y` to save it. ```sh python -m pip install "robot-inference-client[keyring] @ https://yellowandblack.dev/sdk/robot_inference_client-0.5.3-py3-none-any.whl" yb setup --url "https://yellowandblack.dev" ``` If a supported OS keyring is available, setup offers to save your key there. Otherwise, follow its instructions to set `YB_API_KEY` in your shell. Use this installation's address for later CLI commands. In bash/zsh: ```sh export YB_URL="https://yellowandblack.dev" ``` In PowerShell: ```powershell $env:YB_URL = "https://yellowandblack.dev" ``` 3. See which models are offered: ```sh yb models --ready ``` An offered model may be idle. Opening a session starts it; listing models does not start one or reserve capacity. ## Send your first request The SDK includes a recorded LIBERO simulator observation. No separate download or source checkout is needed. Send it to an offered `pi05-libero` model: ```python from robot_inference_client import Client, example_observation task, observation = example_observation() with Client(base_url="https://yellowandblack.dev") as client, client.session( model="pi05-libero", instruction=task, max_spend_usd="1" ) as session: result = session.infer(observation) print(result["actions"].shape) # (10, 7): the next 10 actions ``` Save it as `first_call.py`, and run it: ```sh python first_call.py ``` Opening the session waits for the model to become ready. Add `on_status=print` to `client.session(...)` to display `starting` and `ready` when those states are reached. Startup and waiting are free; they are separate from inference time. If the SDK's startup waiting limit expires, it raises `model_start_timeout` and requests cancellation. By default it uses the model's advertised waiting limit. You can set a shorter limit with `connect_timeout`; this does not make the model start faster. See [Session options](https://yellowandblack.dev/docs/python-sdk.md#session-options) for checking whether cancellation was confirmed before opening another session. > **Note:** `max_spend_usd="1"` caps the session at $1 of available credit. This is replay > input: its send time is not a camera capture time, and its output does not demonstrate > task success. The recorded example uses the LIBERO schema, not DROID. The same example is available from the terminal: ```sh yb run --model pi05-libero --example --requests 1 --max-spend-usd 1 ``` For your own recorded LIBERO observation, use `--fixture observation.npz` instead of `--example`. A failed inference preserves its JSON report and exits nonzero; temporary retryable failures use exit status 75. ### Test your setup for free `transport-demo` is a free sandbox. It isn't a model: it returns placeholder actions in the same format, so you can check your key, network and code without spending credit. ```python from robot_inference_client import Client from robot_inference_client.cli import synthetic_observation with Client(base_url="https://yellowandblack.dev") as client, client.session( model="transport-demo", instruction="test" ) as session: result = session.infer(synthetic_observation(0, "test")) print(result["actions"].shape) # (10, 7): placeholder actions ``` Or check it from a terminal. This sends 10 placeholder observations, then prints how long each answer took and what the session cost: ```sh yb run --model transport-demo ``` For a DROID robot, use `yb run --model transport-demo-droid`. ## Run it in a control loop Replace the placeholder images with data from your robot or simulator, and call `infer` once per step: ```python from robot_inference_client import Client, InferenceError task = "put the bowl on the plate" with Client(base_url="https://yellowandblack.dev") as client, client.session( model="pi05-libero", instruction=task, max_spend_usd="5", transport="owned", # bounds the SDK's socket waits ) as session: while robot.running(): captured = session.capture_time_ns() # when the cameras take the pictures observation = { "observation/image": robot.main_camera(), # uint8, (224, 224, 3) "observation/wrist_image": robot.wrist_camera(), # uint8, (224, 224, 3) "observation/state": robot.state(), # 8 numbers: arm position and gripper "prompt": task, } try: result = session.infer(observation, captured_ns=captured) except InferenceError as error: if error.retry == "next_request": continue # this answer came too late or was replaced: send the next one raise robot.execute(result["actions"]) # the next moves ``` Each model expects its own data format; [Robot data formats](https://yellowandblack.dev/docs/robot-contracts.md) lists them. ## Next steps - [Python SDK](https://yellowandblack.dev/docs/python-sdk.md): every class and method, and each session option - [Billing](https://yellowandblack.dev/docs/billing.md): prices, credit and spending caps - [Errors](https://yellowandblack.dev/docs/errors.md): what each error means, and how to fix it --- # Guide for AI agents Learn how an AI coding agent should read these docs, get a key, handle errors, and check its work. ## Read the docs as text - [/llms.txt](https://yellowandblack.dev/llms.txt) lists every page, with a one-line summary each. - [/llms-full.txt](https://yellowandblack.dev/llms-full.txt) holds every page in one file. - Any page is markdown when you add `.md` to its address, for example [/docs/quickstart.md](https://yellowandblack.dev/docs/quickstart.md). Asking for `text/markdown` in the `Accept` header works too. - In a terminal, `yb docs` lists the pages and `yb docs ` prints one, for example `yb docs robot-contracts`. No key is needed. - [/openapi.json](https://yellowandblack.dev/openapi.json) describes every HTTP route. ## Get a key A person must do two things in the dashboard first: 1. Create an account. Verify its email before adding credit; the free sandbox doesn't need it. Accept the service terms too, if the dashboard asks. 2. Create an API key in the **API keys** tab. Then give the agent the key as `YB_API_KEY`. > **Note:** Never put the key in code, in a command-line argument, in a committed file, > or in a chat. ## Start safely - Start with `transport-demo` and `max_spend_usd="0"`. It's a free sandbox, not a model: it returns placeholder actions, and it can't spend money. - Write `max_spend_usd` as a string, such as `"5"`, never as a float. - On a real robot, open sessions with `transport="owned"`, so a stalled network can't freeze the control loop. - Open sessions with `with client.session(...) as policy:`, so they always close. An open session holds a place on the model, and its reserved credit, until it closes or times out. - If your code may retry `client.session(...)` after a crash, pass the same `idempotency_key` each time. The same key never opens a second session. - Never send generated test data to a real model: its answers would be meaningless. Use recorded data, or the robot itself. ## Handle errors Every error has a `code`, such as `capacity`. The Python SDK and `yb` also give a `fix`, a `docs_url`, and a `retry` value. Plain HTTP responses give only `code` and `message`, sometimes with a true or false `retry`; see [HTTP API](https://yellowandblack.dev/docs/http-api.md#errors). Use `retry` to decide what to do next: | `retry` | What the agent should do | |---|---| | `retry` | Wait, then repeat the same call. Wait longer after each failure, and give up after a few tries. For a `PlatformError`, stop at once if `error.retryable` is `False`. | | `fix` | Do what the `fix` text says first. The same call again fails the same way. | | `next_request` | Only this request failed, and the session is still open. Send the next observation. | | `new_session` | The session has ended. Open a new one if the task still needs it. | | `none` | A normal ending. Nothing to do. | For example: ```python from robot_inference_client import InferenceError, PlatformError captured = policy.capture_time_ns() # read when the cameras take the pictures try: result = policy.infer(observation, captured_ns=captured) except InferenceError as error: if error.retry != "next_request": raise # the session ended; error.fix says what to do # the answer was late or replaced: skip it and send the next observation except PlatformError as error: print(error.code, error.fix, error.docs_url) raise ``` `yb errors ` explains any code without going online. `yb` exits with status 75 for a temporary error, 1 for other errors, and 2 for a wrong command line. ## Check the integration 1. `yb models --ready` lists the offered model. If it is on demand and idle, the session waits for startup; otherwise check that capacity is available. Listing alone does not start a model. 2. Each observation matches the model's format on [Robot data formats](https://yellowandblack.dev/docs/robot-contracts.md): the exact names, `uint8` images of shape (224, 224, 3), the right number of state values, and a `prompt` equal to the session's task. 3. `captured_ns` comes from `policy.capture_time_ns()`, read when the cameras take the pictures. 4. The loop treats errors with `retry == "next_request"` as skipped steps, not as crashes. 5. Every session closes on every path, including when an exception is raised. 6. After closing, `policy.closure` isn't `None`, `policy.closure["close_confirmed"]` is `True`, and `policy.closure["charged_usd"]` is what you expect. --- # Python SDK reference The Python SDK lets your code open sessions on our models, send your robot's data, and read what you spent. It requires Python 3.8 or newer. Use a virtual environment as shown in [Quickstart](https://yellowandblack.dev/docs/quickstart.md), then install it with: ```sh python -m pip install "https://yellowandblack.dev/sdk/robot_inference_client-0.5.3-py3-none-any.whl" ``` You import it as `robot_inference_client`. It has two main classes: - `Client` connects to the platform. You use it to list models, open sessions, and check your usage. - `PolicySession` is one open session. You use it to send observations and get the next moves. ## Client `example_observation(model="pi05-libero")` is also exported by the SDK. It returns `(instruction, observation)` from the bundled recorded LIBERO frame, with independent arrays on each call. It supports `pi05-libero` and `transport-demo`; no download or source checkout is needed. This replay input has no live capture timestamp and does not qualify task success. See [Quickstart](https://yellowandblack.dev/docs/quickstart.md) for a complete call. ```python Client(api_key=None, base_url=None, project_id=None) ``` With no arguments, `Client` connects to the hosted service with the key saved by `yb setup`. `YB_API_KEY` overrides that key. It checks the key when you create it. A wrong key raises `AuthorizationError` with the code `unauthorized`. Use it in a `with` block, or call `close()` when you're done. Closing a `Client` doesn't close your sessions: close each one. | Method | What it does | |---|---| | `models(ready=False)` | With `ready=True`, offered models: `model_id`, `status`, `mode` (`warm` or `on_demand`), `rate`, session slots and `contract`. On-demand entries may be idle until a connection starts. Reading this reserves nothing. Without it, every catalog model is returned. | | `session(model, instruction, max_spend_usd="0", label="Python client", max_action_age_ms=2000, mode="client", idempotency_key=None, transport="sync", connect_timeout=None, on_status=None)` | Waits for model readiness, opens the connection and returns a `PolicySession`. See the options below. | | `sessions()` | Your project's sessions: state, completed requests, cost, and why each one ended. | | `model_session(session_id)` | One session, with the model server's live counts while it's open. | | `close_session(session_id)` | Closes a session by its ID. Safe to call twice. | | `usage()` | Your credit balance, the credit reserved by open sessions, and your ledger. | | `regions(model)` and `select_region(model, region="auto")` | For operators only: measure the network delay to each region. | | `deploy(...)`, `deployment(deployment_id)` and `deployments()` | For operators and developers only: a private model server that one project owns. You don't need these to use the running models. | | `close()` | Closes the client's connection to the platform. | ### Session options - `model`: a `model_id` from `models(ready=True)`. - `connect_timeout`: maximum seconds waiting for on-demand model startup; defaults to the model's advertised startup limit. This is separate from individual HTTP and worker handshake timeouts. Temporary `database_busy` or `executor_busy` status responses are polled again within this same limit. Admission and inference requests are not replayed. A readiness response received at or after the startup limit is treated as `model_start_timeout`; the SDK requests cancellation without renewing a grant. - `on_status`: optional callback receiving `"starting"` and `"ready"`, for example `on_status=print`. Loading and waiting are included; they are not inference charges. Ctrl-C or a startup timeout requests cancellation. A startup `PlatformError` includes `session_id`, `cleanup_requested` and `cleanup_confirmed` in its context; if requested cleanup was not confirmed, call `close_session` or check its status. Server-side startup and attach deadlines still expire the admission. - `instruction`: the task in words, such as `"put the bowl on the plate"`. Every observation's `prompt` must match it. Change it with `reset`. - `max_spend_usd`: the most this session may cost, as a string or `Decimal`, such as `"5"`. A float is refused. It's reserved from your credit while the session is open, and it also caps the number of requests. `"0"` works only on the free sandbox. - `max_action_age_ms`: the deadline, meaning how old an answer may be when it reaches you, counted from when the cameras took the pictures. The default is 2000, and the range is 20 to 30000. Older answers are dropped. For example, `max_action_age_ms=500` drops any answer older than half a second. An answer the model server drops as late is free. An answer the server sent in time, but that reached you late, is charged. - `idempotency_key`: if your code retries after a crash, pass the same value each time, so a retry never opens a second session. By default, each call gets a new value. While that session is open, the same value returns it. Once it has ended, you get an error such as `session_closed`: open a new session with a new value. If the admission response is lost, the `PlatformError` context includes the generated `idempotency_key`; use it to recover that admission. A failed connection attempt on an already active replay does not automatically close the other connection. - `transport`: `"sync"` (the default) or `"owned"`. With `"owned"`, every network wait gives up at the request's deadline, so a stuck network can't freeze your robot's loop. Use `"owned"` for a control loop on a real robot. It needs `websockets` 13 or newer. The default is fine for simulators. - `image_codec`: `"jpeg"` (the default) sends camera images as JPEG, quality 90, which is 14 times fewer bytes than raw pixels. π0.5 succeeded on 30 of 30 LIBERO tasks with it, against 29 of 30 with raw pixels (measured). `"raw"` sends exact pixels, for example to compare against a local model bit for bit. A model server that doesn't accept JPEG gets raw pixels automatically. `jpeg_quality` (50 to 100, default 90) sets the quality. - `udp`: `"auto"` (the default), `"on"` or `"off"`. With `"auto"`, observations go over UDP when the model server offers it and a test packet gets through as the session opens; otherwise they go over the WebSocket. UDP doesn't wait for lost packets to be sent again: spare packets rebuild most losses on arrival, and the server asks for the rest. On a link with a 20 ms round trip that lost 1% of packets each way, the slowest 1% of requests took 119 ms over UDP and 356 ms over the WebSocket (measured, including 86 ms of model time). With no loss the two tie. Answers come back both ways, and you get whichever arrives first. If two UDP requests in a row get no answer at all, the session switches to the WebSocket by itself, and tries UDP again every 30 seconds. `"on"` raises `fastpath_unavailable` when UDP can't be used as the session opens; `"off"` never tries UDP. The `YB_UDP` variable sets the default. UDP needs outbound UDP to the model server's port; see [Security](https://yellowandblack.dev/docs/security.md) for how it's encrypted. One robot's open session on one model. Use one per robot, from one thread at a time: its methods must not run at the same time. | Member | What it does | |---|---| | `infer(observation, captured_ns=None)` | Sends one observation (camera images, arm state and task) and waits for the answer. Returns a dict: `actions` (the next moves, one row per step, as a NumPy array), `client_action_age_ms` (how old the answer was on arrival), `sequence` and `epoch` (which request and which episode it answers), `model_revision` (which model version answered) and `timing`. Raises `InferenceError` if the request fails. | | `submit(observation, captured_ns=None)` | Sends one observation without waiting, and returns its `sequence` number. Collect the answer with `result`. Send the next observation while the robot still executes the current actions, so the model's time is hidden behind the motion. The model server keeps only the newest waiting observation per session, so submitting again before the last one started replaces it (`superseded`). | | `result(sequence, timeout=None)` | Waits for the answer to a submitted observation, until its deadline, and returns the same dict as `infer`. With `timeout` (seconds) it waits less and raises `TimeoutError`, leaving the request pending, so your loop can do other work and ask again. | | `capture_time_ns()` | The SDK's clock, in nanoseconds. Read it when the cameras take the pictures, and pass it as `captured_ns`. | | `network()` | How the session reaches the model server: `path` (`"udp"` or `"websocket"`), `udp_round_trip_ms` (measured as the session opened), `note` (why UDP isn't in use, if it isn't), `answers` (how many answers arrived first by each path) and `udp` (packet counts, including resends). | | `reset(instruction=None)` | Starts a new episode (a new attempt at the task), optionally with a new task. Answers meant for the old episode are dropped. | | `status()` | The platform's view of the session: state, completed requests, cost so far, and the model server's counts. | | `renew_credentials()` | Gets a fresh session pass now. You rarely need it: the SDK renews the pass once half its life has passed, and through a short platform outage it keeps trying until the pass is about to expire. | | `report()` | The model server's report on the session: counts and timings. | | `flush_telemetry()` | Sends up to 100 recent timings measured on your side, to help find problems. It never changes what you pay. | | `close()` | Ends the session and settles its cost. Safe to call twice. Leaving a `with` block calls it. | | `closure` | Set by `close()`: `completed` (requests answered), `charged_usd`, `settlement_source` (see [Billing](https://yellowandblack.dev/docs/billing.md)), `close_confirmed` and `close_reason`. `close_reason` is `customer_closed` when your code closed the session, or `worker:` plus the model server's reason when it had already ended the session, such as `worker:too_many_expired`. It's `None` if the platform couldn't be reached; the session then ends on its own, and nothing extra is charged. If the model server couldn't be reached, `closure` shows the state `closing` and `charged_usd` 0 until the platform settles it. | | `assignment` | The platform's record of the session: model, version, the model server it runs on (`generation`), price, spending cap and request limit. | | `id`, `model_id`, `model_revision`, `generation` | Which session this is, which model version it runs, and which model server it's on. | ## Errors Every error has a `code`, such as `capacity`. [Errors](https://yellowandblack.dev/docs/errors.md) says what each code means and how to fix it. | Class | Raised when | Extra members | |---|---|---| | `PlatformError` | a call to the platform fails. Codes not listed below, such as `terms_required` or `email_unverified`, raise a plain `PlatformError`. | `status` (the HTTP status), `context` (extra details), `retryable` | | `CapacityError` | `no_capacity`, `capacity`, `session_limit`, `account_session_limit` or `rate_limit` (a kind of `PlatformError`) | the same | | `AuthorizationError` | `unauthorized`, `forbidden`, `not_found` or `rate_changed` | the same | | `SpendError` | `insufficient_credit`, `spend_too_small` or `spend_limit` | the same | | `InferenceError` | `infer` or `reset` fails: for example a late answer (`expired`), a closed session, or a dropped connection (`connection_lost`). | none | Every error also has `code`, `message`, `fix`, `retry` and `docs_url`. `retry` says what to do next: `retry`, `fix`, `next_request`, `new_session` or `none`. The table on [Errors](https://yellowandblack.dev/docs/errors.md) explains each. Printing an error shows all of these. For example, to retry when the model is full, waiting longer each time: ```python import time from robot_inference_client import CapacityError def open_session(client, instruction, attempts=5): for attempt in range(attempts): try: return client.session("pi05-libero", instruction, max_spend_usd="5") except CapacityError as error: if not error.retryable or attempt == attempts - 1: raise time.sleep(2**attempt) # nothing is charged while you wait ``` ## Environment variables | Variable | Meaning | |---|---| | `YB_URL` | Another copy of the platform to use instead of the hosted service. You don't need it otherwise. | | `YB_API_KEY` | Your key, instead of the one saved by `yb setup`. For servers and CI. | | `YB_PROJECT_ID` | Which project to use, when one key can see more than one. | | `YB_ALLOW_INSECURE_HTTP` | Set to `1` to allow plain `http` to another machine on a private network you control. Without it, the SDK sends keys only over `https`, or to this machine. | | `YB_UDP` | The default for the `udp` session option: `auto`, `on` or `off`. | ## Robot safety The SDK checks that each answer belongs to the request you sent, and that it's fresh when it arrives. It can't know whether the moves still fit the world when your robot makes them. Your robot's controller must check that sensor data is fresh, that each move still makes sense, and when to stop. --- # Command line (`yb`) `yb` is the command-line tool that comes with the Python SDK. You can use it to check your key, list models, run a test session, and look up errors, without writing code. Use `--json` for scripts and agents: results and errors use the versioned format below, and progress stays on standard error. Without it, commands retain their terminal output: most results are JSON, while docs, help, setup instructions and logs use text. ## Scripts and agents Put `--json` before or after the command, for example `yb --json models --ready` or `yb models --ready --json`. A finite command writes one JSON record followed by a newline: ```json {"schema_version":1,"command":"models","type":"result","ok":true,"data":[],"error":null,"context":{"project_id":"example-project"}} ``` `schema_version` is the envelope version, currently 1. `command` is null when argument parsing cannot identify it. `data` holds the normal result or useful partial results; `error` is null on success. `context` contains available project, model, deployment, session, request, generation and idempotency identifiers. Missing identifiers are omitted. Consumers should tolerate additional fields and unknown error codes. On failure, `ok` is false and `error` contains `code`, `message`, `next_step`, `retry`, `source_retry`, `retryable`, `http_status` (possibly null) and `docs`. `source_retry` preserves the error reference's category; `retry` describes the next action in this CLI context. For example, an expired observation normally permits a next request, but `yb run` ends its batch and closes the session, so its advice reflects that closure. Unknown server codes are preserved with conservative, nonretryable guidance. `yb run` retains observed counts, timings, errors and known closure/charge fields even when a batch fails or is interrupted. `requests_sent` counts SDK calls begun, not proof of server receipt or billable completion. Unknown settlement fields stay null. Exit 75 describes a temporary failure; it does not mean it is safe to replay robot actions. For an ambiguous deployment request, inspect its state and reuse the returned `context.idempotency_key` with `--idempotency-key` for the same operation. The CLI does not automatically retry mutations. `yb logs DEPLOYMENT --follow --json` emits newline-delimited JSON (NDJSON): `type:"event"` records followed by one `type:"end"` record on an orderly finish, handled interruption or error while output remains writable. The terminal data includes `last_sequence` and `events_received`. Resume with `--after N` using the last event your consumer processed. A terminal deployment state triggers one final event read; this is not a promise that no later event can ever be written. Reading logs from a failed deployment can succeed: the end record then has `ok:true` and `deployment_status:"failed"`. Without `--follow`, JSON mode returns one result containing an `events` list. Forced process termination or an unwritable output stream cannot provide an end record. Help uses `type:"help"` with `data.text`; docs use `type:"result"` with `data.markdown`. Both remain available without a project key. `yb setup --json` never prompts or writes to a keyring: it uses `YB_API_KEY` or one piped input line, reports verified project/origin and `stored:false`, and fails usefully when input is absent. Human `yb setup` retains its interactive keyring consent. `yb doctor --json` keeps its findings under `data` and returns `diagnostic_findings` with exit 1 when changes are needed. ## Everyday commands | Command | What it does | |---|---| | `yb setup [--url URL] [--no-keyring]` | Asks for your key without showing it, checks it, and can save it in your system keyring. It writes only its own entry, and asks before replacing a saved key. Outside a terminal, it reads the key from standard input. | | `yb models [--ready]` | With `--ready`, offered models and their connection state. An on-demand model may be idle until you connect. Without it, every model in the catalog. | | `yb run [--model M] [--requests N] [--max-spend-usd S]` | Runs a short test session: sends observations, prints each answer's timing and the cost, then closes. | | `yb sessions` | Your project's sessions, with state, completed requests and cost. | | `yb close ` | Closes a session. Safe to repeat. The model keeps running. | | `yb usage` | Your credit balance, credit reserved by open sessions, and your ledger. | | `yb docs [page] [--url URL]` | Lists the docs pages, or prints one as markdown. No key needed. | | `yb errors [code]` | Explains an error code without going online, or lists every code. When a code also names a way a session can end, such as `expired`, it explains both. No key needed. | | `yb doctor [--seconds N] [--url URL]` | Checks what slows your robot's link: Wi-Fi power saving, the Wi-Fi band and signal, and the round trip to your router compared with the service. It prints what to fix, and exits with 1 if it found something. It changes no settings. No key needed. | ## `yb run` options | Option | Default | Meaning | |---|---|---| | `--model` | `transport-demo` | The model to call. The default is the free sandbox; for a DROID robot, use `transport-demo-droid`. | | `--instruction` | `transport test` | The task, in words. | | `--requests` | 10 | How many observations to send. | | `--interval-ms` | 100 | Milliseconds between observations. | | `--budget-ms` | 2000 | The deadline for each answer (`max_action_age_ms`), in milliseconds. | | `--max-spend-usd` | `0` | The most the session may cost. | | `--transport` | `sync` | `owned` makes every network wait give up at the deadline. Use it on a real robot. | | `--fixture` | none | Recorded input: an `.npz` file with `image`, `wrist_image`, `state` and `prompt`. This loader supports the LIBERO schema. Mutually exclusive with `--example`. | | `--example` | off | Replay the recorded LIBERO frame bundled with the SDK; supports `pi05-libero` and `transport-demo`. | With the default, `transport-demo` (a free sandbox that returns placeholder actions), `yb run` generates placeholder data. A learned LIBERO model needs `--fixture` or `--example`; placeholder images are not meaningful task inputs. It also needs `--max-spend-usd`, such as `1`, because the default, `0`, works only on the free sandbox. ## Operator and developer commands These commands manage a private model server that one project owns, started and stopped on request. You don't need them to connect to the offered models. | Command | What it does | |---|---| | `yb deploy [--model M] [--minutes N] [--wait] [--region R]` | Starts a private model server. | | `yb list` | Lists the project's private model servers. | | `yb status ` | Shows one of them. | | `yb stop ` | Stops it. | | `yb logs [--follow] [--after N]` | Prints events after sequence N (default 0); JSON follow mode uses NDJSON. | | `yb diagnose ` | Measures the network delay to it, not the model's time. | | `yb demo ` | Sends one placeholder observation to a test server. | | `yb regions [--model M]` | Measures the network delay to each region. | ## Exit status | Status | Meaning | |---|---| | 0 | Success. | | 1 | An error or actionable diagnostic findings. JSON mode returns the error envelope on stdout; terminal mode uses stderr for errors. | | 2 | The command line was wrong. | | 75 | A temporary error, such as a full model. Retry later, and wait longer each time. | | 130 | The command was interrupted. Inspect any returned operation IDs and partial data before retrying. | ## Environment `yb` reads the same variables as the SDK: `YB_URL`, `YB_API_KEY`, `YB_PROJECT_ID` and `YB_ALLOW_INSECURE_HTTP`. See [Python SDK](https://yellowandblack.dev/docs/python-sdk.md#environment-variables). --- # Robot data formats Each model expects your robot's data in one exact format, and replies in one exact format. This page shows both, for every model: the names to use, the size of each camera image and list of numbers, and the shape of the moves that come back. - The model server checks your data, and refuses data in the wrong format. With the SDK, one observation in the wrong format ends the session with `invalid_observation`. - The model server always replies in the format shown here, and the SDK checks each answer. - `yb models --ready` shows each model's format, under `contract`. Camera images are 224 × 224 color (RGB) arrays of `uint8`. The robot's state is a list of ordinary numbers (no NaN or infinity). The `prompt` is the task in words, and must match the session's task; change it with `reset`. ## `openpi-libero-v1` Robot: LIBERO simulator (Franka, robosuite). How it moves: end-effector delta + gripper, about 10 steps per second. | Send | Type | Shape | Meaning | |---|---|---|---| | `observation/image` | uint8 | 224 × 224 × 3 | Main camera, RGB | | `observation/wrist_image` | uint8 | 224 × 224 × 3 | Wrist camera, RGB | | `observation/state` | float | 8 | eef_x, eef_y, eef_z, axis_angle_x, axis_angle_y, axis_angle_z, gripper_0, gripper_1 | | `prompt` | str | | The session's task, in words | Returns `actions`: an array of numbers of shape (10, 7): one row per step, one column per number. Columns: dx, dy, dz, drot_x, drot_y, drot_z, gripper. Carry out the first few steps, then send a new observation. Recorded September 25 model test (historical): the π0.5 LIBERO model was tested on a rented A100 GPU on 2026-09-25 (86 ms per request). Repeated runs on the same input give slightly different actions (largest difference about 0.004); we are still working to make them identical. For later Modal serving and startup measurements, see [Performance and availability](https://yellowandblack.dev/docs/faq.md#how-fast-is-it). These dated records do not qualify a physical robot or guarantee current capacity. ```python import numpy as np observation = { "observation/image": np.zeros((224, 224, 3), dtype=np.uint8), "observation/wrist_image": np.zeros((224, 224, 3), dtype=np.uint8), "observation/state": np.zeros(8, dtype=np.float32), "prompt": "put the bowl on the plate", } ``` ## `openpi-droid-v1` Robot: DROID (Franka Panda, joint-velocity control). How it moves: joint velocity + gripper position, about 15 steps per second. | Send | Type | Shape | Meaning | |---|---|---|---| | `observation/exterior_image_1_left` | uint8 | 224 × 224 × 3 | Exterior camera, RGB | | `observation/wrist_image_left` | uint8 | 224 × 224 × 3 | Wrist camera, RGB | | `observation/joint_position` | float | 7 | q1, q2, q3, q4, q5, q6, q7 | | `observation/gripper_position` | float | 1 | gripper_position | | `prompt` | str | | The session's task, in words | Returns `actions`: an array of numbers of shape (15, 8): one row per step, one column per number. Columns: dq1, dq2, dq3, dq4, dq5, dq6, dq7, gripper_position. Carry out the first few steps, then send a new observation. Recorded September 25 model test (historical): the π0.5 DROID model was tested on a rented A100 GPU on 2026-09-25 (87 ms per request). It has not yet been tested on a real DROID robot. For later Modal serving and startup measurements, see [Performance and availability](https://yellowandblack.dev/docs/faq.md#how-fast-is-it). These dated records do not qualify a physical robot or guarantee current capacity. ```python import numpy as np observation = { "observation/exterior_image_1_left": np.zeros((224, 224, 3), dtype=np.uint8), "observation/wrist_image_left": np.zeros((224, 224, 3), dtype=np.uint8), "observation/joint_position": np.zeros(7, dtype=np.float32), "observation/gripper_position": np.zeros(1, dtype=np.float32), "prompt": "put the bowl on the plate", } ``` --- # Errors Every error has a short code, such as `capacity`. Find it below to see what happened and what to do. - **Python:** errors have `code`, `fix`, `retry` and `docs_url`, a link to the code's row on this page. - **HTTP:** the body is `{"error": {"code": "...", "message": "..."}}`. - **Terminal:** `yb errors ` explains a code without going online. ## What to do next Each code has a `retry` value. It tells your program what to do next: | `retry` | What to do | |---|---| | `retry` | Temporary: wait, then repeat the same call with backoff. | | `fix` | Change what the fix says first; repeating the same call fails the same way. | | `next_request` | Only this request failed; the session is still open, send the next observation. | | `new_session` | This session is over; open a new session if you still need one. | | `none` | A normal ending; nothing to do. | ## Opening a session | Code | What happened | What to do | `retry` | |---|---|---|---| | `no_capacity` | No ready copy of this model is running right now. It may be warming up, being replaced, or switched off. Nothing was charged. | Retry with backoff. If the response's `retry` flag is false (in Python, `error.retryable` is False), the model is not offered at all; pick one from `yb models --ready`. | `retry` | | `capacity` | The model is running, but every session slot on it is taken. Nothing was charged. | Retry shortly, or close one of your own open sessions. | `retry` | | `session_limit` | This project already has its maximum number of open sessions. | Close a session you no longer use: `yb sessions`, then `yb close `. | `fix` | | `account_session_limit` | The account's open sessions across all its projects reached the ceiling. | Close a session in any of the account's projects. | `fix` | | `rate_changed` | The model's price changed after your client read it. | Read the new price with `yb models --ready`, then open a new session. | `fix` | | `insufficient_credit` | Available credit does not cover `max_spend_usd` (or, for a dedicated deployment, its operating window). | Add credit in the dashboard, or lower `max_spend_usd`. | `fix` | | `spend_limit` | `max_spend_usd` is above the per-session limit. | Use a smaller `max_spend_usd`. Open another session when this one runs out. | `fix` | | `spend_too_small` | `max_spend_usd` does not cover even one request at this model's rate. | Raise `max_spend_usd`. The rate is listed by `yb models --ready`. | `fix` | | `model_unavailable` | No enabled model has that name. | Pick a model from `yb models --ready`. | `fix` | | `idempotency_required` | Opening a session or starting a deployment needs an Idempotency-Key header of 8 to 128 characters. | Send one unique key per logical attempt. The SDK does this for you. | `fix` | | `idempotency_conflict` | This Idempotency-Key was already used with different parameters. | Use a new key for a different request. Reuse a key only to retry the exact same request. | `fix` | | `controller_fenced` | The platform's control service stopped making changes to stay safe, because another copy of it may be running. Nothing was changed or charged. | Retry later. An operator must restart the control service. | `retry` | | `expiring` | The session or deployment is about to reach its end time, so no new session pass is issued. | Open a new session. | `new_session` | | `session_expired` | The session reached its absolute time limit. | Open a new session. | `new_session` | | `session_closed` | The session is already closed. A closed session can never be reopened. | Open a new session. | `new_session` | | `replica_retired` | The model server this session was pinned to was replaced, by a planned renewal or after a failure. Unfinished work was not charged. | Open a new session. It lands on the current model server. | `new_session` | | `health_stale` | The model server has not reported its health recently. | Retry shortly. A model server that stays silent is replaced automatically. | `retry` | | `worker_unhealthy` | The model server reports that it is not healthy. | Retry shortly. An unhealthy model server is replaced automatically. | `retry` | | `not_ready` | The model server or deployment is not accepting sessions yet. | Wait until it is ready, then retry. | `retry` | | `unavailable` | The platform, an upstream service or the model server is temporarily unavailable. | Retry shortly. Check the operation's state before retrying a mutation and reuse its idempotency key. Do not replay robot observations or actions. | `retry` | | `draining` | The model server is shutting down, and takes no new sessions. | Open a new session. The platform sends it to the current copy. | `retry` | | `service_budget` | The platform has no more funded GPU capacity right now. | Retry later. | `retry` | | `deadline_too_short` | The answer deadline (`max_action_age_ms`) is shorter than this model can ever meet, so every answer would arrive too late. | Raise `max_action_age_ms` to at least the model's minimum, listed by `yb models --ready` under `limits`. | `fix` | | `slot_share` | This account already holds its share of one model server's places. The rest are kept for other customers. | Close one of your sessions on this model, or wait for one to end. | `fix` | ## During a session | Code | What happened | What to do | `retry` | |---|---|---|---| | `expired` | A result would have been too old to use, so it was dropped: the observation waited too long, the model finished after the deadline, or the SDK measured capture-to-receipt age above `max_action_age_ms`. Results the model server drops are not charged. A result the server sent in time, but that reached you late, is. | Send the next observation. If it happens often, raise `max_action_age_ms`, share the model with fewer clients, or run closer to its region. | `next_request` | | `superseded` | A newer observation arrived before this one started, so this one was dropped. Only the newest waiting observation is kept. | Nothing to do. To get every result, wait for each one before sending the next observation. | `next_request` | | `invalid_observation` | The observation does not match the model's robot contract: keys, image shape, state size, or prompt. After 20 invalid observations in a row the session closes. | Compare it with the Robot data formats page in the docs. Keep the prompt equal to the session instruction; change it with `reset`. | `fix` | | `sequence` | Sequence numbers must increase within an episode. | Let the SDK number requests. Numbering restarts after `reset`. | `fix` | | `identity` | A request or session pass does not match the session: a different session, episode or model, or the model server opened a session other than the one assigned. | Open a new session through the SDK. Never reuse IDs or session passes across sessions. | `new_session` | | `worker_failed` | The model process failed or timed out, so its sessions closed. Nothing was charged for unfinished work. | Open a new session after a short wait. A replacement copy starts automatically. | `new_session` | | `spend_exhausted` | The session used up the request allowance its `max_spend_usd` paid for. | Open a new session with a new `max_spend_usd`. | `new_session` | | `already_connected` | Another client is already connected to this session. A session has exactly one client. | Use one session per robot or process. Open another session for another client. | `fix` | ## Keys, accounts and permissions | Code | What happened | What to do | `retry` | |---|---|---|---| | `unauthorized` | No valid key or sign-in: the API key is missing, revoked, expired (keys last 90 days) or mistyped, or the dashboard sign-in ended. | Set YB_API_KEY to a current project key from the dashboard's API keys tab, or sign in again. | `fix` | | `forbidden` | The key or sign-in is valid, but may not do this. Changing the account needs a signed-in person, not an API key. | Do this step in the dashboard while signed in. | `fix` | | `scope` | A session pass was used outside its own session. A pass opens exactly one session, on one model server. | Use the SDK's session object for that session, and the project API key for everything else. | `fix` | | `origin` | A browser request came from a website other than this platform's own address. | Call the API from the dashboard itself or from code running outside a browser. | `fix` | | `login_failed` | The email or the password is wrong. | Check both, or reset the password from the sign-in page. | `fix` | | `password_mismatch` | The current password you typed is wrong. | Type the account's current password exactly. | `fix` | | `account_exists` | An account already uses this email. | Sign in instead, or reset that account's password. | `fix` | | `account_changed` | The signed-in account no longer matches the account shown when this action was submitted. | Reload the dashboard, confirm the current account and project, then submit the action again. | `fix` | | `signup_disabled` | This installation is not accepting new accounts. | Ask the operator for access. | `fix` | | `email_unverified` | The account email is not verified yet, and this action can spend money. | Open the verification link from your email, or request a new one in the dashboard. | `fix` | | `email_unavailable` | The verification email could not be sent. | Try again in a few minutes. Contact support if it keeps failing. | `retry` | | `email_unconfigured` | This installation cannot send email, so password reset by email is off. | Ask the operator to reset the password. | `fix` | | `link_expired` | The email link is invalid, expired or already used. | Request a new link. | `fix` | | `terms_required` | The account has not accepted the current service terms. | A person signs in to the dashboard and accepts them. An API key cannot accept terms. | `fix` | | `terms_version_mismatch` | The terms version you accepted is not the one currently published. | Reload the dashboard and accept the version it shows. | `fix` | | `key_limit` | The project already has the maximum number of API keys. | Revoke a key you no longer use, then create a new one. | `fix` | | `project_limit` | The account already has the maximum of five projects. | Use one of the existing projects. | `fix` | | `not_found` | That project, session, deployment or model does not exist, or this key cannot see it. | Check the ID. List what you can see with `yb sessions` or `yb models --ready`. | `fix` | | `rate_limit` | Too many attempts of one kind in a short time: sign-ins, emails, session opens, deployment starts or checkouts. | Wait, then retry with backoff. Reuse one session for many requests instead of opening one per request. | `retry` | | `last_project` | The account must retain at least one project. | Rename this project or create another before removing an unused project. | `fix` | | `project_has_history` | A project with credit, usage or agreements cannot be removed as unused. | Keep its accounting history; contact support if cleanup is needed. | `fix` | | `recovery_hold` | Public service access is temporarily quarantined for recovery verification. | Retry later or contact support; do not create a new account to bypass recovery. | `retry` | ## Checks inside the Python SDK | Code | What happened | What to do | `retry` | |---|---|---|---| | `invalid_argument` | A command argument or local input is invalid. | Check the message and run `yb --help`; correct the input before retrying. | `fix` | | `input_required` | Noninteractive operation needs input that was not supplied. | Set YB_API_KEY or pipe one project key to `yb setup --json`; keys never belong in command-line arguments. | `fix` | | `interrupted` | The command was interrupted. An accepted remote operation may still exist. | Inspect the returned session/deployment IDs and usage before retrying; reuse the idempotency key for the same management mutation. | `fix` | | `diagnostic_findings` | The diagnostic completed and found issues. | Read the findings in the result, make the indicated changes, then run the diagnostic again. | `fix` | | `local_io_error` | A local file or input/output operation failed. | Check that the input file exists and is readable, then retry the command. | `fix` | | `unexpected_error` | The command encountered an unexpected implementation or response error. | Inspect any returned operation IDs and contact support with the SDK version before retrying a mutation. | `fix` | | `unreachable` | The SDK could not reach the platform: network, DNS or TLS failed, or the server is down. | Check YB_URL and your network, then retry with backoff. | `retry` | | `closed` | A method was called on a session that is already closed. | Open a new session. | `new_session` | | `reset` | The session was reset elsewhere, so results for the previous episode were discarded. | Continue with the session's current instruction. | `next_request` | | `reset_timeout` | The server did not confirm a reset in time. | Close the session and open a new one. | `new_session` | | `revision` | The model server runs a different model version than the one admitted, so the SDK sent nothing. | Open a new session. | `new_session` | | `schema` | The model server's data format is not the one this client expects, or this SDK version does not know the format. | Upgrade the SDK, or pick the model whose contract your robot code implements. | `fix` | | `invalid_result` | A result did not belong to the request that was sent (session, episode or sequence). | Open a new session. Report it if it repeats. | `new_session` | | `connection_lost` | The connection to the model server stopped working: it closed, an observation could not be sent in time, or nothing came back for several requests in a row. The SDK closed the session. | Open a new session. If it keeps happening, check the network between the robot and the service. | `new_session` | | `fastpath_unavailable` | You asked for UDP (udp="on") and the session cannot use it: the model server does not offer it, the pycryptodomex package is missing, or UDP to the model server is blocked. | Allow outbound UDP to the model server's Fastpath port, install pycryptodomex, or use udp="auto" (the default), which falls back to the WebSocket by itself. | `fix` | | `invalid_actions` | The model server returned an action chunk with the wrong shape or non-finite values; the SDK refused it. | Open a new session and report the model ID to support. | `new_session` | | `invalid_credential` | A session pass renewal tried to move the session to another model server or model version; the SDK refused it. | Open a new session. | `new_session` | | `model_unknown` | No ready model has that name. | List models with `yb models --ready`. | `fix` | | `request_failed` | The server returned an error without a code. | Read the message and retry once. Report it if it repeats. | `fix` | | `model_start_timeout` | Model startup exceeded the SDK connection deadline. No inference ran through this connection. | Check the error's cleanup_confirmed field or close its session_id, then open a new session. | `new_session` | | `model_start_failed` | The connection closed before model startup completed. Model loading is not a customer charge. | Read close_reason in the error context. Retry with a new session after the cause is resolved. | `new_session` | ## Credit and payments | Code | What happened | What to do | `retry` | |---|---|---|---| | `checkout_unknown` | The payment provider response was lost or unavailable; the purchase outcome is unknown. | Retry with the SAME Idempotency-Key or inspect payment history. Do not start another purchase just to retry. | `retry` | | `topup_range` | The top-up amount is outside the allowed range for one purchase. | Choose an amount between the `min_usd` and `max_usd` in the error. | `fix` | | `payments_unconfigured` | Payments are not connected on this installation. | Ask the operator to add credit. | `fix` | | `payment_mismatch` | A payment event did not match the recorded checkout, so it was refused and nothing changed. | Nothing for customers to do. The operator investigates. | `fix` | | `payment_pending` | The internal purchase or original payment needed to reconcile this event is not recorded yet. | The durable inbox retries it. If it persists, the operator reconciles the missing payment record. | `retry` | | `invalid_signature` | A payment webhook had an invalid signature and was refused. | Nothing for customers to do. The operator checks the webhook secret. | `fix` | | `credit_forfeiture_required` | Deletion would forfeit unused purchased credit. | Review the balance and explicitly acknowledge forfeiture, or keep the account open. | `fix` | | `payment_review` | The provider outcome requires manual reconciliation. | Contact support with the payment ID; do not repay to recover the same purchase. | `fix` | | `review_resolution` | The operator outcome or case note is missing or invalid. | Choose a supported outcome and a short reconciliation reference. | `fix` | | `refund_unconfirmed` | A full refund has not been confirmed by Stripe. | Reconcile the provider event before recording a completed refund. | `fix` | | `payment_paid` | A paid purchase cannot be marked uncharged. | Choose the verified payment outcome; preserve its accounting record. | `fix` | | `payment_unconfirmed` | A purchase has not been confirmed as paid. | Reconcile it with the provider before deciding how its credit is handled. | `fix` | | `credit_unavailable` | Held credit cannot be released to an inactive account or an unconfirmed purchase. | Review account eligibility and the provider payment first. | `fix` | | `account_not_closed` | An active or suspended account's credit cannot be forfeited as a deleted account. | Resolve access and release its confirmed credit, or reconcile an exceptional refund. | `fix` | | `recovery_pending` | A restored account needs review before a payment event can apply. | The operator reconciles account ownership, deletion and balances; the inbox retains the event. | `fix` | ## Dedicated deployments (operators and developers only) | Code | What happened | What to do | `retry` | |---|---|---|---| | `source_resource_unverified` | Cleanup is held because the saved resource belongs to an unverified source installation. | The operator must reconcile source-host ownership and provider inventory before confirming cleanup. | `fix` | | `duration_limit` | The requested deployment duration is longer than the service allows. | Ask for a shorter duration. | `fix` | | `project_capacity` | The project already has an active deployment. | Stop it before starting another. | `fix` | | `not_a_deployment_model` | This model is offered as a shared ready model, not as a dedicated deployment. | Open a session on it instead: `client.session(...)` or `yb run`. | `fix` | | `not_synthetic` | This action needs a ready CPU sandbox deployment. | Use a sandbox deployment for the demo. | `fix` | | `synthetic_worker` | LIBERO simulator runs need a real model backend, not the CPU sandbox. | Run LIBERO against a GPU model server. | `fix` | | `worker_unreachable` | The model server's telemetry could not be read. | Retry shortly. | `retry` | | `busy` | A model server reconnect was requested while sessions are open or the server is still starting. | Close active sessions and wait for start-up to finish, then reconnect. | `fix` | | `region_unavailable` | No enabled model profile exists in the requested region. | Pick another region, or `auto`. | `fix` | | `ambiguous_model` | More than one profile matches that model in the region. | Name an exact model profile. | `fix` | | `probes_unavailable` | Region probes could not be measured from this client. | Name an explicit region instead of `auto`. | `fix` | | `deployment_unavailable` | The deployment failed or is stopping. | Read its events with `yb logs `, then start a new one. | `fix` | | `readiness_timeout` | The deployment did not become ready in time. It still exists and may still be billed by the hour. | Inspect it with `yb status `, or stop it with `yb stop `. | `fix` | ## Generic HTTP errors | Code | What happened | What to do | `retry` | |---|---|---|---| | `database_busy` | The platform's database is busy, closing or unable to finish within its wait limit. | Wait for Retry-After, check the operation's current state, and reuse its idempotency key when retrying a mutation. Do not replay robot observations or actions. | `retry` | | `executor_busy` | The platform's bounded background-processing capacity is temporarily occupied. | Wait for Retry-After and retry with backoff. Reuse the same idempotency key for a management mutation. | `retry` | | `bad_request` | The request was malformed. | Check the method, path and body against the HTTP API page in the docs. | `fix` | | `method_not_allowed` | This path does not accept that HTTP method. | Check the HTTP API page in the docs. | `fix` | | `conflict` | The request conflicts with the current state. | Read the current state, then retry. | `fix` | | `too_large` | The request body is larger than allowed (1 MiB for management calls). | Send a smaller body. | `fix` | | `invalid_request` | A field is missing or has the wrong type or range. The error's `fields` list names each one. | Fix the named fields. | `fix` | ## Why a session ended Every closed session records why it ended, as `close_reason`: in `policy.closure` after `close()`, and in `client.sessions()`. A reason that starts with `worker:`, such as `worker:too_many_expired`, means the model server ended the session before your code closed it. Some reasons share a name with an error code, such as `expired`; this table gives the meaning for a session that ended. During a session, the SDK raises an `InferenceError` with the model server's reason as its code. When the platform ends a session, for example for `credit_reversed`, the model server reports `client_closed`; the session's `close_reason` has the real reason. | Reason | What happened | What to do | `retry` | |---|---|---|---| | `client_closed` | Your client closed the session. | Nothing to do. | `none` | | `startup_timeout` | Model startup exceeded the platform deadline. The admission's credit hold was released. | Check model availability before opening another connection. | `new_session` | | `startup_failed` | The worker stopped before admitting inference. Model loading was not charged. | Wait for the service to recover, then open a new connection. | `new_session` | | `customer_closed` | The session was closed through the API, `yb close`, or the dashboard. | Nothing to do. | `none` | | `disconnected` | The client's connection dropped, so the session closed. | Open a new session. Check the network if it repeats. | `new_session` | | `idle_expired` | No traffic for longer than the idle limit, or no client connected in time. | Open a new session. Close sessions you are not using. | `new_session` | | `attach_expired` | The session was admitted, but no client connected within the attach window (60 seconds by default). Its slot and credit hold were released. | Connect right after opening, or open a new session. | `new_session` | | `credential_expired` | The session's pass expired without being renewed. | Open a new session. The SDK renews automatically between requests. | `new_session` | | `expired` | The session reached its end time. | Open a new session. | `new_session` | | `deployment_expired` | The model server reached its own end time. | Open a new session. | `new_session` | | `replica_retired` | The model server was replaced, by a planned renewal or after a failure. Unfinished work was not charged. | Open a new session. | `new_session` | | `worker_failed` | The model process failed. Unfinished work was not charged. | Open a new session after a short wait. | `new_session` | | `gateway_stopped` | The model server shut down. | Open a new session. | `new_session` | | `gateway_restarted` | The model server restarted while the session was open. | Open a new session. | `new_session` | | `slow_consumer` | The client stopped reading results and 32 unread messages piled up. | Read results as they arrive, then open a new session. | `new_session` | | `too_many_rejections` | Twenty invalid observations arrived in a row. | Fix the observation format (see the Robot data formats page in the docs), then open a new session. | `new_session` | | `too_many_expired` | Ten model runs in a row finished after their deadline, so none could be delivered. | Raise `max_action_age_ms`, send fresher observations, or run closer to our region, then open a new session. | `new_session` | | `credit_reversed` | A refund or dispute took back credit this session's spending cap relied on. Completed requests were settled. | Add credit, then open a new session. | `new_session` | | `account_deleted` | The account was deleted. | Nothing to do. | `none` | | `restored` | The platform was restored from a backup; sessions open at that moment were closed without charge. | Open a new session. | `new_session` | | `test_complete` | A console sandbox test finished. | Nothing to do. | `none` | | `test_error` | A console sandbox test failed. | Run the test again. | `none` | --- # HTTP and WebSocket API Call the API directly when you write a client in another language, or to debug. [/openapi.json](https://yellowandblack.dev/openapi.json) describes the platform HTTP routes: request and response fields, authentication, required headers, response codes, and downloads. The model server and its binary session stream are separate interfaces, described below. Responses may add fields; clients should ignore fields they do not recognize. > **Note:** If you use Python, use the [Python SDK](https://yellowandblack.dev/docs/python-sdk.md). It does everything > on this page, and checks that each answer matches your request and is fresh. ## Two servers - **The platform**, at `https://yellowandblack.dev`, handles keys, credit, and opening and closing sessions. Its routes start with `/v1/platform/`. - **A model server** runs one model. When you open a session, the platform returns the server's address as `endpoint`, and a session pass as `grant`. Your robot's data goes to the assigned inference gateway. In the Modal layout the gateway has its own container on the CPU host and forwards observations to a private GPU worker. ## Authentication Send your key to the platform: ``` Authorization: Bearer yb_live_... ``` Send the session pass to the model server. Never send it your key. ``` Authorization: Bearer ``` A pass works for one session, for about 15 minutes. To renew it, get a new pass with `POST .../model-sessions/{id}/renew`, then send the new pass to the model server with `POST /v1/sessions/{id}/renew`. Otherwise the session closes with `credential_expired` when the old pass runs out. ## Errors Every error has the same shape: ```json { "error": { "code": "capacity", "message": "The ready model has no free session slot; retry shortly", "retry": true } } ``` Look up `code` on [Errors](https://yellowandblack.dev/docs/errors.md). `retry`, when present, says whether repeating the call can help. Some errors also carry `rate` (the model's current price) or `fields` (the inputs that were wrong). ## Safe retries `POST .../model-sessions` needs an `Idempotency-Key` header: a value you choose, 8 to 128 characters. Send the same value when you retry, so a retry never opens a second session. The same value with a different body fails with `idempotency_conflict`. A ready new session returns `201`; an idempotent replay returns `200`. On-demand startup returns `202`, including an idempotent replay that is still starting. It has `state: starting`, `startup_by`, and no `grant`. Poll the session URL; once its state is `granted`, call its `/renew` route to obtain a scoped grant. `DELETE` cancels startup and releases its credit hold. Startup is not billed. The Python SDK handles this wait automatically. Replaying a finished session returns its record without a new grant. Deployment creation and checkout also require an `Idempotency-Key` of 8 to 128 characters. ## Platform routes Paths start with `/v1/platform`. `{project}` is your project's ID, from `GET /me`. | Route | What it does | |---|---| | `GET /meta` | The model catalog and service settings. No key needed. | | `GET /status` | Whether each model is working. No key needed. | | `GET /me` | Your account and its projects. | | `GET /projects/{project}/model-services` | Offered models, including on-demand models, with `mode`, startup limit, price, data format and session slots. Reading this starts no GPU. | | `POST /projects/{project}/model-sessions` | Admits a connection. A ready session includes `endpoint` and `grant`; a starting session returns HTTP 202 without a grant. | | `GET /projects/{project}/model-sessions` | Your sessions. | | `GET /projects/{project}/model-sessions/{id}` | One session, with live counts while it's open. | | `POST /projects/{project}/model-sessions/{id}/renew` | A fresh pass for the same session. | | `DELETE /projects/{project}/model-sessions/{id}` | Closes the session and settles its cost. Safe to repeat. | | `GET /projects/{project}/usage` | Your balance, reserved credit and ledger. | | `GET /projects/{project}/usage.csv` | The ledger, as CSV. | Exports are live reports, not frozen snapshots. Concurrent changes may be absent or reflect different times. Account JSON's `exported_at` is the generation start. Exports cannot be resumed. Usage CSV succeeds only after the report is fully prepared, up to 8 MiB. To verify receipt, require `X-YB-Export-Version: 1`, then compare the **decoded UTF-8 response bytes** with `X-YB-Export-Bytes` (decimal byte count) and `X-YB-Export-SHA256` (lowercase hexadecimal SHA-256). Do not use HTTP `Content-Length` for decoded-byte verification when compression is present. `X-YB-Export-Rows` counts CSV records excluding the header; quoted fields may contain newlines. The console performs the byte-count and digest checks before offering a download. Keep the metadata if you need to verify a saved CSV later. `export_too_large` returns 422; `export_busy`, `export_timeout` and `export_failed` return 503. These generation errors return JSON instead of CSV. A transfer failure after headers can only interrupt the response: reject any incomplete or unverifiable body. Account JSON does not use these CSV headers. The body for opening a session: | Field | Meaning | |---|---| | `model_id` | A model from `model-services`. | | `instruction` | The task in words, such as `"put the bowl on the plate"`. | | `max_spend_usd` | The most the session may cost, as a string, such as `"5"`. The default, `"0"`, works only on the free sandbox. | | `rate_version` | The `rate.version` you read from `model-services`. Needed for paid models. If the price changed since, the call fails with `rate_changed`. | | `max_action_age_ms` | Optional. How old an answer may be when it reaches you. The default is 2000. | | `label` | Optional. A name for the session in the dashboard. | | `mode` | Optional. Keep the default, `client`. | ## Model server routes Paths are relative to the session's `endpoint`. | Route | What it does | |---|---| | `POST /v1/sessions` | Joins the session the pass names. Body: `instruction`, `label`, `mode`, `max_action_age_ms`. Check that the response names the same session ID and model version as the platform's record. Stop if it doesn't. | | `GET /v1/sessions/{id}/stream` | Opens the WebSocket that carries observations and answers. One client per session. | | `POST /v1/sessions/{id}/renew` | Moves the session to a new pass. Send the new pass as the bearer token. | | `POST /v1/sessions/{id}/reset` | Starts a new episode, optionally with a new `instruction`. | | `DELETE /v1/sessions/{id}` | Closes the session on the model server. | ## The session stream Each session has one WebSocket. Every message is a binary [MessagePack](https://msgpack.org) map. NumPy arrays travel as maps with the keys `__ndarray__`, `dtype`, `shape` and `data`. Each key is a MessagePack binary string, not a text string. `dtype` is a NumPy type string such as `|u1`, and `data` holds the array's bytes in C order. Turn WebSocket compression off. The SDK's `codec.py` and `client.py` are the reference implementation. 1. The model server first sends `{"type": "ready", ...}`, naming the session. 2. You send one observation: ```python { "version": 1, "type": "observation", "session_id": "ses_...", "epoch": 0, # the episode: goes up by one with each reset "sequence": 1, # goes up by one with each observation in an episode "captured_ns": 0, # your clock when the cameras took the pictures; sent back to you "budget_ms": 2000.0, # time left before the answer is useless "observation": {...}, # the model's data format } ``` 3. The server answers `{"type": "actions", "actions": , ...}`: the next moves. The answer repeats `session_id`, `epoch`, `sequence`, `captured_ns` and `model_revision`. Check each one against your request before using the moves. 4. Or it answers `{"type": "error", "code": ..., "sequence": ...}`, for example `expired` (too late) or `superseded` (replaced by newer data). The session stays open. 5. `{"type": "reset", ...}` means answers for earlier episodes no longer count. `{"type": "closed", "reason": ...}` ends the session. [Errors](https://yellowandblack.dev/docs/errors.md#why-a-session-ended) lists the reasons. --- # Billing and credit You pay per request the model answers. You pay from prepaid credit, and each session has a spending cap that you set. ## Prices Each model has a price per 1,000 requests. The **Models** tab shows it, and so does `yb models --ready`, under `rate`. - `transport-demo` is free. It's a sandbox for testing your setup, not a model: it returns placeholder actions. - Using a model means you accept its price. There's nothing to click first. - Check the published rate before opening a paid session. Open sessions keep their original rate. Supplying `rate_version` rejects a new session if its rate has changed. ## How a session is billed 1. **Start with trial credit.** The launch offer gives each eligible new user **$5** in promotional credit once, automatically after email verification and acceptance of the current service terms. There is no total giveaway cap. Deleting an account and registering again with the same email does not grant another trial. To purchase more credit, open **Usage & billing**; the configured purchase range is displayed there (default $25–$100). Promotional credit is spent first. Unused promotional and purchased credits are nonrefundable and are forfeited when you confirm account deletion. 2. **Open a session with a cap.** `max_spend_usd` is the most the session can cost. That amount is reserved from your credit while the session is open. 3. **The cap limits requests.** A session can make at most its cap divided by the price of one request, rounded down. At that point it stops with `spend_exhausted`. Answers the model server drops as late aren't charged, but they still count toward this limit, because the model ran for them. 4. **Close the session.** You pay the requests the model answered, times the price, never more than the cap. The unused reservation is released when settlement completes. An unreachable worker can leave the session closing while cleanup retries. Each session is settled exactly once. The free sandbox needs no credit, and has no request limit. An on-demand model may show **Starting** while capacity is prepared. Your spending cap is held during this wait, but no inference is charged. Cancelling or timing out before admission releases the hold. Other customers using the shared model keep their connections. After the last connection ends, we may shut down its GPU. ## Example For example, at a made-up price of $2.00 per 1,000 requests: - `max_spend_usd="5"` reserves $5.00 and allows up to 2,500 requests. - A robot that sends ten requests a second for three minutes sends 1,800 requests. If the model answers all of them, the session costs $3.60. - When the session closes, $3.60 is charged and $1.40 comes back to your credit. ## What's free - Opening a session, loading a model, starting it up, and idle time. - Requests refused because of their format, the key, a full model, or a rate limit. - Answers the model server drops because they would be too late (`expired`), because newer data replaced them (`superseded`), or because the model failed (`worker_failed`). An answer the server sent in time, but that reached your code late, is still charged. - Unfinished work when a model server is replaced (`replica_retired`). ## Refunds, disputes and lost servers - A refunded or disputed card payment removes that credit at once. If the remaining credit no longer covers what open sessions reserved, they close with `credit_reversed`, and are settled for what they completed. - A released dispute hold restores only the corresponding held credit. Settled refunds remain removed. - If the machine running a model server is lost, its open sessions are charged the last count the server reported, never more than their caps. With no report, they're charged nothing. ## See your usage - `yb usage`, or `client.usage()`, shows your balance, reserved credit and ledger. - In the **Usage & billing** tab, **Export CSV** downloads the ledger. The console checks that the complete generated file arrived before offering the download. The report is generated live, so concurrent changes may not be included. Retry a reported verification failure; contact support if the report exceeds the 8 MiB download limit. - **Purchases** shows pending, completed, expired, and review states. Leaving the checkout page does not confirm or cancel a payment. Resume its link or cancel the unfinished purchase from this history. - Checkout retries reuse the same purchase identifier. If a response is lost, check the existing purchase before starting a new one. A payment ID marked for review lets support reconcile an uncertain provider outcome. - Account deletion disables access first and continues cleanup in the background. Keep the deletion receipt to check progress. Existing unused credits are forfeited after settlement; a payment that completes during deletion is held for separate reconciliation and never restores account access. - A closed session shows `completed` (requests answered), `charged_usd` and `settlement_source`, which says how the cost was settled: - `worker`: by the model server's own count - `snapshot`: by the model server's final count, saved when the server was replaced - `last_report`: the machine was lost, so the server's last reported count was used - `unavailable`: nothing was known, so nothing was charged > **Note:** A charged request means our server sent the answer. It doesn't prove that your > robot ran it, or that it still fit when it did. Timings your code reports help find > problems, but the model server's count settles the bill. --- # Limits These are the limits the service enforces today, and why each one exists. They aren't performance promises. The operator can change the values marked "default". ## Sessions | Limit | Value | Why | |---|---|---| | Open sessions per project | 2 (default) | A session takes a place on a shared model server, not a whole GPU. | | Open sessions per account | 4 (default) | The same, across all of an account's projects. | | Places per model server | shown as `capacity` in `yb models --ready` | A few sessions take turns on one server. | | Your share of one model server | all but one of its places (default) | Someone else can always get a place. Going over fails with `slot_share`. | | Session length | 240 minutes (default) | For longer work, open a new session. | | Time to connect after opening | 60 seconds (default) | An unused session gives back its place and its reserved credit. | | Connected but idle | 5 minutes without an observation | The session closes with `idle_expired`, and frees its place. Open a new one when the robot resumes. | | Session pass | lasts about 15 minutes; the SDK renews it at half-life | A leaked pass stops working soon. A short platform outage doesn't end your session: the SDK keeps retrying until the pass is about to expire. | | Sessions opened per project | 120 per hour | Reuse one session for many requests. | | Spending cap per session | $1,000 (default) | Limits what one mistake can cost. | ## Requests | Limit | Value | Why | |---|---|---| | Message size | 8 MiB | One LIBERO observation is about 300 KB (measured). | | Array size | 2,097,152 values, up to 4 dimensions | Oversized arrays are refused before they use memory. | | Deadline (`max_action_age_ms`) | 20 to 30,000 ms; 2,000 by default; at least the model's minimum | An answer older than this is dropped. Answers the model server drops as late aren't charged. A deadline below the model's minimum (120 ms for π0.5, 20 ms for the `transport-demo` sandbox) is refused with `deadline_too_short`, because no answer could arrive in time. | | Late answers in a row | 10 | Then the session closes with `too_many_expired`: the deadline is too short for the model. | | Model runs per session | at most the session's request limit | Answers dropped as late aren't charged, but they count toward the limit, because the model ran for them. | | Waiting observations per session | 1 | A newer observation replaces one that hasn't started. | | Unread answers per session | 32 | Code that stops reading loses its session, instead of slowing everyone down. | | Invalid observations in a row | 20 | Then the session closes with `too_many_rejections`. | | Task length | 1,000 characters | Tasks are short, such as "put the bowl on the plate". | | Request body sent to the platform | 1 MiB | Platform requests are small. | ## Accounts | Limit | Value | |---|---| | Projects per account | 5 | | Key lifetime | 90 days; create a new key before the old one expires | | Sign-in and account attempts | 30 per hour per network address, and 10 per 15 minutes per email | | Verification emails | 3 per hour | | Credit purchases started | 10 per hour per project, and 10 per hour per account (default) | | One credit purchase | $25 to $100 (default) | ## Not limited - The number of moves in one answer. You pay per request, not per move. - Idle time in an open session is never charged. A connected session that sends nothing for 5 minutes still closes, so its place is free for someone else. --- # Security and data How your keys and data are protected, what we keep, and how to report a problem. ## Keys and session passes - **Keys** let code act for one project. We store only a hash of each key, so no one can show it to you again after it's created. You can revoke a key in the **API keys** tab at any time. - **`yb setup`** verifies a key entered at a hidden interactive prompt. When the optional `keyring` dependency and an allowed OS store are available, it offers to save the key for that platform URL after your confirmation and checks that it can read it back. It writes only this service's entry and does not delete other saved passwords. Native credential-store integrations have not yet been qualified. - **Headless setup:** the SDK and subsequent commands can use `YB_API_KEY`. `yb setup --json` reads that variable or one key piped on standard input; it never prompts or saves the key. Without `--json`, noninteractive setup reads one line from standard input. Interactive `--no-keyring` skips the storage offer. Setup does not write a plaintext credential file or fall back to file-based keyrings. - **Session passes** let the SDK use one session on one model server, for about 15 minutes. The API calls a pass a `grant`. A pass can't read, change or close any other session, and it stops working when its server is replaced. The SDK gets and renews passes for you. - **Only people can accept terms or change the account.** A key can't. A leaked key can still spend credit, up to each session's cap, so revoke it quickly. ## Data in transit - Camera images go to the worker address assigned to your session. For Modal-backed models, a Yellow and Black gateway handles the connection and forwards observations to a private Modal GPU function. The production layout uses a separate gateway container for each model generation. The development sandbox runs inside the platform. - The hosted service uses only HTTPS and secure WebSockets. A copy you run on your own machine for development may use plain `http`. - The SDK won't send a key over plain `http`, except to your own machine, or when you set `YB_ALLOW_INSECURE_HTTP=1` for a private network you control. - Before it sends anything, the SDK checks that the model server runs the exact model version the platform promised. - When a session uses UDP (see `udp` in the [Python SDK](https://yellowandblack.dev/docs/python-sdk.md)), each observation and each answer is encrypted with ChaCha20-Poly1305, the cipher TLS 1.3 and WireGuard use. The key is new for every session and reaches your code only over the session's secure WebSocket. Every packet also carries a check value made with that key. The model server checks it first, and silently drops packets that fail it, repeat an old packet, or name an unknown session. ## Data we keep - **Camera images and robot state:** our gateway does not deliberately save observation payloads. For Modal-backed inference, Modal Function inputs and outputs can be retained by Modal for up to seven days; see [Modal's data retention policy](https://modal.com/docs/guide/security#data-retention). Account deletion does not erase these provider copies immediately. Any separately enabled diagnostic capture must be disclosed before use. - **Session records** keep counts, timings, costs, and why each session ended. The task text and label of a closed session are scheduled for redaction after 30 days. Worker reports and copied snapshots are included; unavailable workers can delay confirmation. - **Money records**, such as payments and the ledger, are kept for accounting. - **Trial eligibility** keeps a minimal keyed identity record after deletion to prevent repeated promotional claims. Raw reset/verification links are not stored. - **Export and deletion:** in the **Project settings** tab, **Export my data** downloads your stored account and usage history, without credentials or internal support notes. Account export requires a dashboard login. Deletion disables access immediately and continues cleanup if you close the tab. Keep the deletion receipt to check whether cleanup is pending or complete. Unused nonrefundable credit requires acknowledgment; financial and minimal security records remain. Encrypted backups age out under the deployed backup retention policy; restoring one requires security/deletion review before account access can be restored. Dashboard sign-out does not stop a running robot session. Revoke a compromised key and close affected sessions separately. Worker grants have bounded lifetimes; network cleanup is not a substitute for the robot's local safety controls. A lost key cannot be shown again: revoke it and create another. ## Model integrity Each model runs from checkpoint files, which hold its learned weights, pinned by a hash of their bytes. A model server refuses to start if its files don't match, so a session never runs a silently changed model. ## Isolation - In the Docker deployment, each real model server uses a pinned image, an unprivileged container user and a private network. A separate launcher starts those containers. - In the Modal deployment, the platform invokes a private, version-pinned GPU class using server-side Modal credentials. Customers receive scoped Yellow and Black session passes. Several customers can share that model through the gateway; this is not a dedicated GPU or operating-system sandbox for each customer. - Model servers answer on their own address (`w.` in front of our domain), apart from the dashboard, so their answers never share the dashboard's cookies or storage. - The dashboard loads its fonts from our own server, so opening it sends nothing to a font provider. ## Report a problem Email the contact in [/.well-known/security.txt](https://yellowandblack.dev/.well-known/security.txt). Include the steps to reproduce the problem. Don't access other customers' data or disrupt the service while testing. --- # Questions and answers ## What is Yellow and Black? A service that runs AI models for robots on our GPUs. Today it serves vision-language-action (VLA) models, such as π0.5 from Physical Intelligence. You call them from a real robot or a simulation, and you don't rent, set up or load a GPU. ## What is a session? One robot's connection to one model, from when you open it to when you close it. A session isn't a whole GPU: a few sessions share one model server and take turns. Each session's pass works only for that session. ## What do I pay for? Each request the model answers. Failed requests are free, and so are loading and idle time. You set a spending cap for each session. See [Billing](https://yellowandblack.dev/docs/billing.md). ## Do I need a GPU? No. The model runs on our GPUs. Your robot's computer needs Python 3.8 or newer and an internet connection. ## Which robots and models are supported? `yb models --ready` lists the offered models and their current state. An on-demand model may be idle; opening a session waits for startup before inference can begin. Today that is π0.5, in two robot setups: - **LIBERO**: a simulated Franka robot arm with two cameras. It sends 8 numbers for the arm's position and gripper, and gets back 7 numbers per move. - **DROID**: a Franka robot arm steered by joint speeds. Its model has run on our GPUs, but hasn't been tested with a real robot yet. To test your code for free, use `transport-demo`: a sandbox, not a model, that returns placeholder actions. See [Robot data formats](https://yellowandblack.dev/docs/robot-contracts.md). ## How fast is it? In private tests on October 1, 2026, L40S GPUs served two simultaneous clients per model. Median SDK response time across 500 shared requests per model was **169 ms for LIBERO** and **172 ms for DROID**. The caller was in a US-East datacenter; this does not measure the network between your robot and the service. Shared capacity and your connection affect response time. Opening an idle model is a separate wait. In 15 cold-start trials per model, successful connections had medians of **12 seconds for LIBERO** and **15 seconds for DROID**. The slowest successful connections took 110 and 97 seconds, and one LIBERO trial timed out. These results do not establish a startup guarantee; capacity delays remain unresolved. Startup and waiting are free. Pass `on_status=print` to `client.session(...)` to see connection progress. See [Session options](https://yellowandblack.dev/docs/python-sdk.md#session-options) for the waiting limit and startup-error handling. ## What if my network is slow or drops? The model server drops an answer that would be too late, and doesn't charge for it. An answer it sent in time that reached you late is still charged. Either way, your code gets an `expired` error and sends the next observation. If the connection drops, the SDK raises `connection_lost` and closes the session: open a new one. Your robot's controller decides what the robot does while it waits. ## Do you store my camera images? Our gateway does not deliberately save camera images or robot state. For Modal-backed inference, Modal may retain function inputs and outputs for up to seven days, even when diagnostic capture is disabled. Account deletion does not immediately erase those provider copies. We keep session counts, timings and costs. See [Security and data](https://yellowandblack.dev/docs/security.md) and [Modal's retention policy](https://modal.com/docs/guide/security#data-retention). ## Can AI coding agents use it? Yes. Every docs page is also available as markdown, [/llms.txt](https://yellowandblack.dev/llms.txt) lists them all, and every error has a code. Through the SDK and `yb`, each error also comes with a fix and a link. See [Guide for AI agents](https://yellowandblack.dev/docs/agents.md). ## Can I train or fine-tune a model here? Not yet. You can use the models that `yb models --ready` lists. ## Is it available now? This release is being prepared for public launch. Public signup and hosted model availability have not been verified. On a configured private installation, the free `transport-demo` sandbox lets you check your integration with placeholder actions; it does not qualify a real model or robot task.