Yellow and Black

Docs Reference

HTTP and WebSocket API

Call the API directly when you write a client in another language, or to debug. /openapi.json describes the platform HTTP routes: request and response fields, authentication, required headers, response codes, and downloads. The model server and its binary session stream are separate interfaces, described below. Responses may add fields; clients should ignore fields they do not recognize.

Note: If you use Python, use the Python SDK. It does everything on this page, and checks that each answer matches your request and is fresh.

Two servers

  • The platform, at https://yellowandblack.dev, handles keys, credit, and opening and closing sessions. Its routes start with /v1/platform/.
  • A model server runs one model. When you open a session, the platform returns the server's address as endpoint, and a session pass as grant. Your robot's data goes to the assigned inference gateway. In the Modal layout the gateway has its own container on the CPU host and forwards observations to a private GPU worker.

Authentication

Send your key to the platform:

Authorization: Bearer yb_live_...

Send the session pass to the model server. Never send it your key.

Authorization: Bearer <grant.token>

A pass works for one session, for about 15 minutes. To renew it, get a new pass with POST .../model-sessions/{id}/renew, then send the new pass to the model server with POST /v1/sessions/{id}/renew. Otherwise the session closes with credential_expired when the old pass runs out.

Errors

Every error has the same shape:

{
  "error": {
    "code": "capacity",
    "message": "The ready model has no free session slot; retry shortly",
    "retry": true
  }
}

Look up code on Errors. retry, when present, says whether repeating the call can help. Some errors also carry rate (the model's current price) or fields (the inputs that were wrong).

Safe retries

POST .../model-sessions needs an Idempotency-Key header: a value you choose, 8 to 128 characters. Send the same value when you retry, so a retry never opens a second session. The same value with a different body fails with idempotency_conflict. A ready new session returns 201; an idempotent replay returns 200. On-demand startup returns 202, including an idempotent replay that is still starting. It has state: starting, startup_by, and no grant. Poll the session URL; once its state is granted, call its /renew route to obtain a scoped grant. DELETE cancels startup and releases its credit hold. Startup is not billed. The Python SDK handles this wait automatically. Replaying a finished session returns its record without a new grant. Deployment creation and checkout also require an Idempotency-Key of 8 to 128 characters.

Platform routes

Paths start with /v1/platform. {project} is your project's ID, from GET /me.

Route What it does
GET /meta The model catalog and service settings. No key needed.
GET /status Whether each model is working. No key needed.
GET /me Your account and its projects.
GET /projects/{project}/model-services Offered models, including on-demand models, with mode, startup limit, price, data format and session slots. Reading this starts no GPU.
POST /projects/{project}/model-sessions Admits a connection. A ready session includes endpoint and grant; a starting session returns HTTP 202 without a grant.
GET /projects/{project}/model-sessions Your sessions.
GET /projects/{project}/model-sessions/{id} One session, with live counts while it's open.
POST /projects/{project}/model-sessions/{id}/renew A fresh pass for the same session.
DELETE /projects/{project}/model-sessions/{id} Closes the session and settles its cost. Safe to repeat.
GET /projects/{project}/usage Your balance, reserved credit and ledger.
GET /projects/{project}/usage.csv The ledger, as CSV.

Exports are live reports, not frozen snapshots. Concurrent changes may be absent or reflect different times. Account JSON's exported_at is the generation start. Exports cannot be resumed. Usage CSV succeeds only after the report is fully prepared, up to 8 MiB. To verify receipt, require X-YB-Export-Version: 1, then compare the decoded UTF-8 response bytes with X-YB-Export-Bytes (decimal byte count) and X-YB-Export-SHA256 (lowercase hexadecimal SHA-256). Do not use HTTP Content-Length for decoded-byte verification when compression is present. X-YB-Export-Rows counts CSV records excluding the header; quoted fields may contain newlines. The console performs the byte-count and digest checks before offering a download. Keep the metadata if you need to verify a saved CSV later.

export_too_large returns 422; export_busy, export_timeout and export_failed return 503. These generation errors return JSON instead of CSV. A transfer failure after headers can only interrupt the response: reject any incomplete or unverifiable body. Account JSON does not use these CSV headers.

The body for opening a session:

Field Meaning
model_id A model from model-services.
instruction The task in words, such as "put the bowl on the plate".
max_spend_usd The most the session may cost, as a string, such as "5". The default, "0", works only on the free sandbox.
rate_version The rate.version you read from model-services. Needed for paid models. If the price changed since, the call fails with rate_changed.
max_action_age_ms Optional. How old an answer may be when it reaches you. The default is 2000.
label Optional. A name for the session in the dashboard.
mode Optional. Keep the default, client.

Model server routes

Paths are relative to the session's endpoint.

Route What it does
POST /v1/sessions Joins the session the pass names. Body: instruction, label, mode, max_action_age_ms. Check that the response names the same session ID and model version as the platform's record. Stop if it doesn't.
GET /v1/sessions/{id}/stream Opens the WebSocket that carries observations and answers. One client per session.
POST /v1/sessions/{id}/renew Moves the session to a new pass. Send the new pass as the bearer token.
POST /v1/sessions/{id}/reset Starts a new episode, optionally with a new instruction.
DELETE /v1/sessions/{id} Closes the session on the model server.

The session stream

Each session has one WebSocket. Every message is a binary MessagePack map. NumPy arrays travel as maps with the keys __ndarray__, dtype, shape and data. Each key is a MessagePack binary string, not a text string. dtype is a NumPy type string such as |u1, and data holds the array's bytes in C order. Turn WebSocket compression off. The SDK's codec.py and client.py are the reference implementation.

  1. The model server first sends {"type": "ready", ...}, naming the session.
  2. You send one observation:
{
    "version": 1,
    "type": "observation",
    "session_id": "ses_...",
    "epoch": 0,  # the episode: goes up by one with each reset
    "sequence": 1,  # goes up by one with each observation in an episode
    "captured_ns": 0,  # your clock when the cameras took the pictures; sent back to you
    "budget_ms": 2000.0,  # time left before the answer is useless
    "observation": {...},  # the model's data format
}
  1. The server answers {"type": "actions", "actions": <array>, ...}: the next moves. The answer repeats session_id, epoch, sequence, captured_ns and model_revision. Check each one against your request before using the moves.
  2. Or it answers {"type": "error", "code": ..., "sequence": ...}, for example expired (too late) or superseded (replaced by newer data). The session stays open.
  3. {"type": "reset", ...} means answers for earlier episodes no longer count. {"type": "closed", "reason": ...} ends the session. Errors lists the reasons.

Updated 2026-09-25 · View as markdown