Yellow and Black

Docs Reference

Python SDK reference

The Python SDK lets your code open sessions on our models, send your robot's data, and read what you spent. It requires Python 3.8 or newer. Use a virtual environment as shown in Quickstart, then install it with:

python -m pip install "https://yellowandblack.dev/sdk/robot_inference_client-0.5.3-py3-none-any.whl"

You import it as robot_inference_client. It has two main classes:

  • Client connects to the platform. You use it to list models, open sessions, and check your usage.
  • PolicySession is one open session. You use it to send observations and get the next moves.

Client

example_observation(model="pi05-libero") is also exported by the SDK. It returns (instruction, observation) from the bundled recorded LIBERO frame, with independent arrays on each call. It supports pi05-libero and transport-demo; no download or source checkout is needed. This replay input has no live capture timestamp and does not qualify task success. See Quickstart for a complete call.

Client(api_key=None, base_url=None, project_id=None)

With no arguments, Client connects to the hosted service with the key saved by yb setup. YB_API_KEY overrides that key.

It checks the key when you create it. A wrong key raises AuthorizationError with the code unauthorized.

Use it in a with block, or call close() when you're done. Closing a Client doesn't close your sessions: close each one.

Method What it does
models(ready=False) With ready=True, offered models: model_id, status, mode (warm or on_demand), rate, session slots and contract. On-demand entries may be idle until a connection starts. Reading this reserves nothing. Without it, every catalog model is returned.
session(model, instruction, max_spend_usd="0", label="Python client", max_action_age_ms=2000, mode="client", idempotency_key=None, transport="sync", connect_timeout=None, on_status=None) Waits for model readiness, opens the connection and returns a PolicySession. See the options below.
sessions() Your project's sessions: state, completed requests, cost, and why each one ended.
model_session(session_id) One session, with the model server's live counts while it's open.
close_session(session_id) Closes a session by its ID. Safe to call twice.
usage() Your credit balance, the credit reserved by open sessions, and your ledger.
regions(model) and select_region(model, region="auto") For operators only: measure the network delay to each region.
deploy(...), deployment(deployment_id) and deployments() For operators and developers only: a private model server that one project owns. You don't need these to use the running models.
close() Closes the client's connection to the platform.

Session options

  • model: a model_id from models(ready=True).
  • connect_timeout: maximum seconds waiting for on-demand model startup; defaults to the model's advertised startup limit. This is separate from individual HTTP and worker handshake timeouts. Temporary database_busy or executor_busy status responses are polled again within this same limit. Admission and inference requests are not replayed. A readiness response received at or after the startup limit is treated as model_start_timeout; the SDK requests cancellation without renewing a grant.
  • on_status: optional callback receiving "starting" and "ready", for example on_status=print. Loading and waiting are included; they are not inference charges. Ctrl-C or a startup timeout requests cancellation. A startup PlatformError includes session_id, cleanup_requested and cleanup_confirmed in its context; if requested cleanup was not confirmed, call close_session or check its status. Server-side startup and attach deadlines still expire the admission.
  • instruction: the task in words, such as "put the bowl on the plate". Every observation's prompt must match it. Change it with reset.
  • max_spend_usd: the most this session may cost, as a string or Decimal, such as "5". A float is refused. It's reserved from your credit while the session is open, and it also caps the number of requests. "0" works only on the free sandbox.
  • max_action_age_ms: the deadline, meaning how old an answer may be when it reaches you, counted from when the cameras took the pictures. The default is 2000, and the range is 20 to 30000. Older answers are dropped. For example, max_action_age_ms=500 drops any answer older than half a second. An answer the model server drops as late is free. An answer the server sent in time, but that reached you late, is charged.
  • idempotency_key: if your code retries after a crash, pass the same value each time, so a retry never opens a second session. By default, each call gets a new value. While that session is open, the same value returns it. Once it has ended, you get an error such as session_closed: open a new session with a new value. If the admission response is lost, the PlatformError context includes the generated idempotency_key; use it to recover that admission. A failed connection attempt on an already active replay does not automatically close the other connection.
  • transport: "sync" (the default) or "owned". With "owned", every network wait gives up at the request's deadline, so a stuck network can't freeze your robot's loop. Use "owned" for a control loop on a real robot. It needs websockets 13 or newer. The default is fine for simulators.
  • image_codec: "jpeg" (the default) sends camera images as JPEG, quality 90, which is 14 times fewer bytes than raw pixels. π0.5 succeeded on 30 of 30 LIBERO tasks with it, against 29 of 30 with raw pixels (measured). "raw" sends exact pixels, for example to compare against a local model bit for bit. A model server that doesn't accept JPEG gets raw pixels automatically. jpeg_quality (50 to 100, default 90) sets the quality.
  • udp: "auto" (the default), "on" or "off". With "auto", observations go over UDP when the model server offers it and a test packet gets through as the session opens; otherwise they go over the WebSocket. UDP doesn't wait for lost packets to be sent again: spare packets rebuild most losses on arrival, and the server asks for the rest. On a link with a 20 ms round trip that lost 1% of packets each way, the slowest 1% of requests took 119 ms over UDP and 356 ms over the WebSocket (measured, including 86 ms of model time). With no loss the two tie. Answers come back both ways, and you get whichever arrives first. If two UDP requests in a row get no answer at all, the session switches to the WebSocket by itself, and tries UDP again every 30 seconds. "on" raises fastpath_unavailable when UDP can't be used as the session opens; "off" never tries UDP. The YB_UDP variable sets the default. UDP needs outbound UDP to the model server's port; see Security for how it's encrypted.

One robot's open session on one model. Use one per robot, from one thread at a time: its methods must not run at the same time.

Member What it does
infer(observation, captured_ns=None) Sends one observation (camera images, arm state and task) and waits for the answer. Returns a dict: actions (the next moves, one row per step, as a NumPy array), client_action_age_ms (how old the answer was on arrival), sequence and epoch (which request and which episode it answers), model_revision (which model version answered) and timing. Raises InferenceError if the request fails.
submit(observation, captured_ns=None) Sends one observation without waiting, and returns its sequence number. Collect the answer with result. Send the next observation while the robot still executes the current actions, so the model's time is hidden behind the motion. The model server keeps only the newest waiting observation per session, so submitting again before the last one started replaces it (superseded).
result(sequence, timeout=None) Waits for the answer to a submitted observation, until its deadline, and returns the same dict as infer. With timeout (seconds) it waits less and raises TimeoutError, leaving the request pending, so your loop can do other work and ask again.
capture_time_ns() The SDK's clock, in nanoseconds. Read it when the cameras take the pictures, and pass it as captured_ns.
network() How the session reaches the model server: path ("udp" or "websocket"), udp_round_trip_ms (measured as the session opened), note (why UDP isn't in use, if it isn't), answers (how many answers arrived first by each path) and udp (packet counts, including resends).
reset(instruction=None) Starts a new episode (a new attempt at the task), optionally with a new task. Answers meant for the old episode are dropped.
status() The platform's view of the session: state, completed requests, cost so far, and the model server's counts.
renew_credentials() Gets a fresh session pass now. You rarely need it: the SDK renews the pass once half its life has passed, and through a short platform outage it keeps trying until the pass is about to expire.
report() The model server's report on the session: counts and timings.
flush_telemetry() Sends up to 100 recent timings measured on your side, to help find problems. It never changes what you pay.
close() Ends the session and settles its cost. Safe to call twice. Leaving a with block calls it.
closure Set by close(): completed (requests answered), charged_usd, settlement_source (see Billing), close_confirmed and close_reason. close_reason is customer_closed when your code closed the session, or worker: plus the model server's reason when it had already ended the session, such as worker:too_many_expired. It's None if the platform couldn't be reached; the session then ends on its own, and nothing extra is charged. If the model server couldn't be reached, closure shows the state closing and charged_usd 0 until the platform settles it.
assignment The platform's record of the session: model, version, the model server it runs on (generation), price, spending cap and request limit.
id, model_id, model_revision, generation Which session this is, which model version it runs, and which model server it's on.

Errors

Every error has a code, such as capacity. Errors says what each code means and how to fix it.

Class Raised when Extra members
PlatformError a call to the platform fails. Codes not listed below, such as terms_required or email_unverified, raise a plain PlatformError. status (the HTTP status), context (extra details), retryable
CapacityError no_capacity, capacity, session_limit, account_session_limit or rate_limit (a kind of PlatformError) the same
AuthorizationError unauthorized, forbidden, not_found or rate_changed the same
SpendError insufficient_credit, spend_too_small or spend_limit the same
InferenceError infer or reset fails: for example a late answer (expired), a closed session, or a dropped connection (connection_lost). none

Every error also has code, message, fix, retry and docs_url. retry says what to do next: retry, fix, next_request, new_session or none. The table on Errors explains each. Printing an error shows all of these.

For example, to retry when the model is full, waiting longer each time:

import time

from robot_inference_client import CapacityError


def open_session(client, instruction, attempts=5):
    for attempt in range(attempts):
        try:
            return client.session("pi05-libero", instruction, max_spend_usd="5")
        except CapacityError as error:
            if not error.retryable or attempt == attempts - 1:
                raise
            time.sleep(2**attempt)  # nothing is charged while you wait

Environment variables

Variable Meaning
YB_URL Another copy of the platform to use instead of the hosted service. You don't need it otherwise.
YB_API_KEY Your key, instead of the one saved by yb setup. For servers and CI.
YB_PROJECT_ID Which project to use, when one key can see more than one.
YB_ALLOW_INSECURE_HTTP Set to 1 to allow plain http to another machine on a private network you control. Without it, the SDK sends keys only over https, or to this machine.
YB_UDP The default for the udp session option: auto, on or off.

Robot safety

The SDK checks that each answer belongs to the request you sent, and that it's fresh when it arrives. It can't know whether the moves still fit the world when your robot makes them. Your robot's controller must check that sensor data is fresh, that each move still makes sense, and when to stop.

Updated 2026-09-28 · View as markdown