HTTP and WebSocket API
Call the API directly when you write a client in another language, or to debug. /openapi.json describes the platform HTTP routes: request and response fields, authentication, required headers, response codes, and downloads. The model server and its binary session stream are separate interfaces, described below. Responses may add fields; clients should ignore fields they do not recognize.
Note: If you use Python, use the Python SDK. It does everything on this page, and checks that each answer matches your request and is fresh.
Two servers
- The platform, at
https://yellowandblack.dev, handles keys, credit, and opening and closing sessions. Its routes start with/v1/platform/. - A model server runs one model. When you open a session, the platform returns the
server's address as
endpoint, and a session pass asgrant. Your robot's data goes to the assigned inference gateway. In the Modal layout the gateway has its own container on the CPU host and forwards observations to a private GPU worker.
Authentication
Send your key to the platform:
Authorization: Bearer yb_live_...
Send the session pass to the model server. Never send it your key.
Authorization: Bearer <grant.token>
A pass works for one session, for about 15 minutes. To renew it, get a new pass with
POST .../model-sessions/{id}/renew, then send the new pass to the model server with
POST /v1/sessions/{id}/renew. Otherwise the session closes with credential_expired
when the old pass runs out.
Errors
Every error has the same shape:
{
"error": {
"code": "capacity",
"message": "The ready model has no free session slot; retry shortly",
"retry": true
}
}
Look up code on Errors. retry, when present, says whether repeating
the call can help. Some errors also carry rate (the model's current price) or fields
(the inputs that were wrong).
Safe retries
POST .../model-sessions needs an Idempotency-Key header: a value you choose, 8 to 128
characters. Send the same value when you retry, so a retry never opens a second session.
The same value with a different body fails with idempotency_conflict.
A ready new session returns 201; an idempotent replay returns 200. On-demand
startup returns 202, including an idempotent replay that is still starting.
It has state: starting, startup_by, and no grant. Poll the session URL;
once its state is granted, call its /renew route to obtain a scoped grant.
DELETE cancels startup and releases its credit hold. Startup is not billed.
The Python SDK handles this wait automatically. Replaying a
finished session returns its record without a new grant. Deployment creation and
checkout also require an Idempotency-Key of 8 to 128 characters.
Platform routes
Paths start with /v1/platform. {project} is your project's ID, from GET /me.
| Route | What it does |
|---|---|
GET /meta |
The model catalog and service settings. No key needed. |
GET /status |
Whether each model is working. No key needed. |
GET /me |
Your account and its projects. |
GET /projects/{project}/model-services |
Offered models, including on-demand models, with mode, startup limit, price, data format and session slots. Reading this starts no GPU. |
POST /projects/{project}/model-sessions |
Admits a connection. A ready session includes endpoint and grant; a starting session returns HTTP 202 without a grant. |
GET /projects/{project}/model-sessions |
Your sessions. |
GET /projects/{project}/model-sessions/{id} |
One session, with live counts while it's open. |
POST /projects/{project}/model-sessions/{id}/renew |
A fresh pass for the same session. |
DELETE /projects/{project}/model-sessions/{id} |
Closes the session and settles its cost. Safe to repeat. |
GET /projects/{project}/usage |
Your balance, reserved credit and ledger. |
GET /projects/{project}/usage.csv |
The ledger, as CSV. |
Exports are live reports, not frozen snapshots. Concurrent changes may be absent
or reflect different times. Account JSON's exported_at is the generation start.
Exports cannot be resumed. Usage CSV succeeds only after the report is fully
prepared, up to 8 MiB. To verify receipt, require X-YB-Export-Version: 1, then
compare the decoded UTF-8 response bytes with X-YB-Export-Bytes (decimal
byte count) and X-YB-Export-SHA256 (lowercase hexadecimal SHA-256). Do not use
HTTP Content-Length for decoded-byte verification when compression is present.
X-YB-Export-Rows counts CSV records excluding the header; quoted fields may
contain newlines. The console performs the byte-count and digest checks before
offering a download. Keep the metadata if you need to verify a saved CSV later.
export_too_large returns 422; export_busy, export_timeout and
export_failed return 503. These generation errors return JSON instead of CSV.
A transfer failure after headers can only interrupt the response: reject any
incomplete or unverifiable body. Account JSON does not use these CSV headers.
The body for opening a session:
| Field | Meaning |
|---|---|
model_id |
A model from model-services. |
instruction |
The task in words, such as "put the bowl on the plate". |
max_spend_usd |
The most the session may cost, as a string, such as "5". The default, "0", works only on the free sandbox. |
rate_version |
The rate.version you read from model-services. Needed for paid models. If the price changed since, the call fails with rate_changed. |
max_action_age_ms |
Optional. How old an answer may be when it reaches you. The default is 2000. |
label |
Optional. A name for the session in the dashboard. |
mode |
Optional. Keep the default, client. |
Model server routes
Paths are relative to the session's endpoint.
| Route | What it does |
|---|---|
POST /v1/sessions |
Joins the session the pass names. Body: instruction, label, mode, max_action_age_ms. Check that the response names the same session ID and model version as the platform's record. Stop if it doesn't. |
GET /v1/sessions/{id}/stream |
Opens the WebSocket that carries observations and answers. One client per session. |
POST /v1/sessions/{id}/renew |
Moves the session to a new pass. Send the new pass as the bearer token. |
POST /v1/sessions/{id}/reset |
Starts a new episode, optionally with a new instruction. |
DELETE /v1/sessions/{id} |
Closes the session on the model server. |
The session stream
Each session has one WebSocket. Every message is a binary
MessagePack map. NumPy arrays travel as maps with the keys
__ndarray__, dtype, shape and data. Each key is a MessagePack binary string, not a
text string. dtype is a NumPy type string such as |u1, and data holds the array's
bytes in C order. Turn WebSocket compression off. The SDK's
codec.py and client.py are the reference implementation.
- The model server first sends
{"type": "ready", ...}, naming the session. - You send one observation:
{
"version": 1,
"type": "observation",
"session_id": "ses_...",
"epoch": 0, # the episode: goes up by one with each reset
"sequence": 1, # goes up by one with each observation in an episode
"captured_ns": 0, # your clock when the cameras took the pictures; sent back to you
"budget_ms": 2000.0, # time left before the answer is useless
"observation": {...}, # the model's data format
}
- The server answers
{"type": "actions", "actions": <array>, ...}: the next moves. The answer repeatssession_id,epoch,sequence,captured_nsandmodel_revision. Check each one against your request before using the moves. - Or it answers
{"type": "error", "code": ..., "sequence": ...}, for exampleexpired(too late) orsuperseded(replaced by newer data). The session stays open. {"type": "reset", ...}means answers for earlier episodes no longer count.{"type": "closed", "reason": ...}ends the session. Errors lists the reasons.