# HTTP and WebSocket API

Call the API directly when you write a client in another language, or to debug.
[/openapi.json](https://yellowandblack.dev/openapi.json) describes the platform HTTP routes: request and
response fields, authentication, required headers, response codes, and downloads.
The model server and its binary session stream are separate interfaces, described below.
Responses may add fields; clients should ignore fields they do not recognize.

> **Note:** If you use Python, use the [Python SDK](https://yellowandblack.dev/docs/python-sdk.md). It does everything
> on this page, and checks that each answer matches your request and is fresh.

## Two servers

- **The platform**, at `https://yellowandblack.dev`, handles keys, credit, and opening and closing
  sessions. Its routes start with `/v1/platform/`.
- **A model server** runs one model. When you open a session, the platform returns the
  server's address as `endpoint`, and a session pass as `grant`. Your robot's data goes
  to the assigned inference gateway. In the Modal layout the gateway has its own
  container on the CPU host and forwards observations to a private GPU worker.

## Authentication

Send your key to the platform:

```
Authorization: Bearer yb_live_...
```

Send the session pass to the model server. Never send it your key.

```
Authorization: Bearer <grant.token>
```

A pass works for one session, for about 15 minutes. To renew it, get a new pass with
`POST .../model-sessions/{id}/renew`, then send the new pass to the model server with
`POST /v1/sessions/{id}/renew`. Otherwise the session closes with `credential_expired`
when the old pass runs out.

## Errors

Every error has the same shape:

```json
{
  "error": {
    "code": "capacity",
    "message": "The ready model has no free session slot; retry shortly",
    "retry": true
  }
}
```

Look up `code` on [Errors](https://yellowandblack.dev/docs/errors.md). `retry`, when present, says whether repeating
the call can help. Some errors also carry `rate` (the model's current price) or `fields`
(the inputs that were wrong).

## Safe retries

`POST .../model-sessions` needs an `Idempotency-Key` header: a value you choose, 8 to 128
characters. Send the same value when you retry, so a retry never opens a second session.
The same value with a different body fails with `idempotency_conflict`.
A ready new session returns `201`; an idempotent replay returns `200`. On-demand
startup returns `202`, including an idempotent replay that is still starting.
It has `state: starting`, `startup_by`, and no `grant`. Poll the session URL;
once its state is `granted`, call its `/renew` route to obtain a scoped grant.
`DELETE` cancels startup and releases its credit hold. Startup is not billed.
The Python SDK handles this wait automatically. Replaying a
finished session returns its record without a new grant. Deployment creation and
checkout also require an `Idempotency-Key` of 8 to 128 characters.

## Platform routes

Paths start with `/v1/platform`. `{project}` is your project's ID, from `GET /me`.

| Route | What it does |
|---|---|
| `GET /meta` | The model catalog and service settings. No key needed. |
| `GET /status` | Whether each model is working. No key needed. |
| `GET /me` | Your account and its projects. |
| `GET /projects/{project}/model-services` | Offered models, including on-demand models, with `mode`, startup limit, price, data format and session slots. Reading this starts no GPU. |
| `POST /projects/{project}/model-sessions` | Admits a connection. A ready session includes `endpoint` and `grant`; a starting session returns HTTP 202 without a grant. |
| `GET /projects/{project}/model-sessions` | Your sessions. |
| `GET /projects/{project}/model-sessions/{id}` | One session, with live counts while it's open. |
| `POST /projects/{project}/model-sessions/{id}/renew` | A fresh pass for the same session. |
| `DELETE /projects/{project}/model-sessions/{id}` | Closes the session and settles its cost. Safe to repeat. |
| `GET /projects/{project}/usage` | Your balance, reserved credit and ledger. |
| `GET /projects/{project}/usage.csv` | The ledger, as CSV. |

Exports are live reports, not frozen snapshots. Concurrent changes may be absent
or reflect different times. Account JSON's `exported_at` is the generation start.
Exports cannot be resumed. Usage CSV succeeds only after the report is fully
prepared, up to 8 MiB. To verify receipt, require `X-YB-Export-Version: 1`, then
compare the **decoded UTF-8 response bytes** with `X-YB-Export-Bytes` (decimal
byte count) and `X-YB-Export-SHA256` (lowercase hexadecimal SHA-256). Do not use
HTTP `Content-Length` for decoded-byte verification when compression is present.
`X-YB-Export-Rows` counts CSV records excluding the header; quoted fields may
contain newlines. The console performs the byte-count and digest checks before
offering a download. Keep the metadata if you need to verify a saved CSV later.

`export_too_large` returns 422; `export_busy`, `export_timeout` and
`export_failed` return 503. These generation errors return JSON instead of CSV.
A transfer failure after headers can only interrupt the response: reject any
incomplete or unverifiable body. Account JSON does not use these CSV headers.

The body for opening a session:

| Field | Meaning |
|---|---|
| `model_id` | A model from `model-services`. |
| `instruction` | The task in words, such as `"put the bowl on the plate"`. |
| `max_spend_usd` | The most the session may cost, as a string, such as `"5"`. The default, `"0"`, works only on the free sandbox. |
| `rate_version` | The `rate.version` you read from `model-services`. Needed for paid models. If the price changed since, the call fails with `rate_changed`. |
| `max_action_age_ms` | Optional. How old an answer may be when it reaches you. The default is 2000. |
| `label` | Optional. A name for the session in the dashboard. |
| `mode` | Optional. Keep the default, `client`. |

## Model server routes

Paths are relative to the session's `endpoint`.

| Route | What it does |
|---|---|
| `POST /v1/sessions` | Joins the session the pass names. Body: `instruction`, `label`, `mode`, `max_action_age_ms`. Check that the response names the same session ID and model version as the platform's record. Stop if it doesn't. |
| `GET /v1/sessions/{id}/stream` | Opens the WebSocket that carries observations and answers. One client per session. |
| `POST /v1/sessions/{id}/renew` | Moves the session to a new pass. Send the new pass as the bearer token. |
| `POST /v1/sessions/{id}/reset` | Starts a new episode, optionally with a new `instruction`. |
| `DELETE /v1/sessions/{id}` | Closes the session on the model server. |

## The session stream

Each session has one WebSocket. Every message is a binary
[MessagePack](https://msgpack.org) map. NumPy arrays travel as maps with the keys
`__ndarray__`, `dtype`, `shape` and `data`. Each key is a MessagePack binary string, not a
text string. `dtype` is a NumPy type string such as `|u1`, and `data` holds the array's
bytes in C order. Turn WebSocket compression off. The SDK's
`codec.py` and `client.py` are the reference implementation.

1. The model server first sends `{"type": "ready", ...}`, naming the session.
2. You send one observation:

```python
{
    "version": 1,
    "type": "observation",
    "session_id": "ses_...",
    "epoch": 0,  # the episode: goes up by one with each reset
    "sequence": 1,  # goes up by one with each observation in an episode
    "captured_ns": 0,  # your clock when the cameras took the pictures; sent back to you
    "budget_ms": 2000.0,  # time left before the answer is useless
    "observation": {...},  # the model's data format
}
```

3. The server answers `{"type": "actions", "actions": <array>, ...}`: the next moves. The
   answer repeats `session_id`, `epoch`, `sequence`, `captured_ns` and `model_revision`.
   Check each one against your request before using the moves.
4. Or it answers `{"type": "error", "code": ..., "sequence": ...}`, for example `expired`
   (too late) or `superseded` (replaced by newer data). The session stays open.
5. `{"type": "reset", ...}` means answers for earlier episodes no longer count.
   `{"type": "closed", "reason": ...}` ends the session.
   [Errors](https://yellowandblack.dev/docs/errors.md#why-a-session-ended) lists the reasons.
