Python SDK reference
The Python SDK lets your code open sessions on our models, send your robot's data, and read what you spent. It requires Python 3.8 or newer. Use a virtual environment as shown in Quickstart, then install it with:
python -m pip install "https://yellowandblack.dev/sdk/robot_inference_client-0.5.3-py3-none-any.whl"
You import it as robot_inference_client. It has two main classes:
Clientconnects to the platform. You use it to list models, open sessions, and check your usage.PolicySessionis one open session. You use it to send observations and get the next moves.
Client
example_observation(model="pi05-libero") is also exported by the SDK. It returns
(instruction, observation) from the bundled recorded LIBERO frame, with independent
arrays on each call. It supports pi05-libero and transport-demo; no download or source
checkout is needed. This replay input has no live capture timestamp and does not qualify
task success. See Quickstart for a complete call.
Client(api_key=None, base_url=None, project_id=None)
With no arguments, Client connects to the hosted service with the key saved by
yb setup. YB_API_KEY overrides that key.
It checks the key when you create it. A wrong key raises AuthorizationError with the
code unauthorized.
Use it in a with block, or call close() when you're done. Closing a Client doesn't
close your sessions: close each one.
| Method | What it does |
|---|---|
models(ready=False) |
With ready=True, offered models: model_id, status, mode (warm or on_demand), rate, session slots and contract. On-demand entries may be idle until a connection starts. Reading this reserves nothing. Without it, every catalog model is returned. |
session(model, instruction, max_spend_usd="0", label="Python client", max_action_age_ms=2000, mode="client", idempotency_key=None, transport="sync", connect_timeout=None, on_status=None) |
Waits for model readiness, opens the connection and returns a PolicySession. See the options below. |
sessions() |
Your project's sessions: state, completed requests, cost, and why each one ended. |
model_session(session_id) |
One session, with the model server's live counts while it's open. |
close_session(session_id) |
Closes a session by its ID. Safe to call twice. |
usage() |
Your credit balance, the credit reserved by open sessions, and your ledger. |
regions(model) and select_region(model, region="auto") |
For operators only: measure the network delay to each region. |
deploy(...), deployment(deployment_id) and deployments() |
For operators and developers only: a private model server that one project owns. You don't need these to use the running models. |
close() |
Closes the client's connection to the platform. |
Session options
model: amodel_idfrommodels(ready=True).connect_timeout: maximum seconds waiting for on-demand model startup; defaults to the model's advertised startup limit. This is separate from individual HTTP and worker handshake timeouts. Temporarydatabase_busyorexecutor_busystatus responses are polled again within this same limit. Admission and inference requests are not replayed. A readiness response received at or after the startup limit is treated asmodel_start_timeout; the SDK requests cancellation without renewing a grant.on_status: optional callback receiving"starting"and"ready", for exampleon_status=print. Loading and waiting are included; they are not inference charges. Ctrl-C or a startup timeout requests cancellation. A startupPlatformErrorincludessession_id,cleanup_requestedandcleanup_confirmedin its context; if requested cleanup was not confirmed, callclose_sessionor check its status. Server-side startup and attach deadlines still expire the admission.instruction: the task in words, such as"put the bowl on the plate". Every observation'spromptmust match it. Change it withreset.max_spend_usd: the most this session may cost, as a string orDecimal, such as"5". A float is refused. It's reserved from your credit while the session is open, and it also caps the number of requests."0"works only on the free sandbox.max_action_age_ms: the deadline, meaning how old an answer may be when it reaches you, counted from when the cameras took the pictures. The default is 2000, and the range is 20 to 30000. Older answers are dropped. For example,max_action_age_ms=500drops any answer older than half a second. An answer the model server drops as late is free. An answer the server sent in time, but that reached you late, is charged.idempotency_key: if your code retries after a crash, pass the same value each time, so a retry never opens a second session. By default, each call gets a new value. While that session is open, the same value returns it. Once it has ended, you get an error such assession_closed: open a new session with a new value. If the admission response is lost, thePlatformErrorcontext includes the generatedidempotency_key; use it to recover that admission. A failed connection attempt on an already active replay does not automatically close the other connection.transport:"sync"(the default) or"owned". With"owned", every network wait gives up at the request's deadline, so a stuck network can't freeze your robot's loop. Use"owned"for a control loop on a real robot. It needswebsockets13 or newer. The default is fine for simulators.image_codec:"jpeg"(the default) sends camera images as JPEG, quality 90, which is 14 times fewer bytes than raw pixels. π0.5 succeeded on 30 of 30 LIBERO tasks with it, against 29 of 30 with raw pixels (measured)."raw"sends exact pixels, for example to compare against a local model bit for bit. A model server that doesn't accept JPEG gets raw pixels automatically.jpeg_quality(50 to 100, default 90) sets the quality.udp:"auto"(the default),"on"or"off". With"auto", observations go over UDP when the model server offers it and a test packet gets through as the session opens; otherwise they go over the WebSocket. UDP doesn't wait for lost packets to be sent again: spare packets rebuild most losses on arrival, and the server asks for the rest. On a link with a 20 ms round trip that lost 1% of packets each way, the slowest 1% of requests took 119 ms over UDP and 356 ms over the WebSocket (measured, including 86 ms of model time). With no loss the two tie. Answers come back both ways, and you get whichever arrives first. If two UDP requests in a row get no answer at all, the session switches to the WebSocket by itself, and tries UDP again every 30 seconds."on"raisesfastpath_unavailablewhen UDP can't be used as the session opens;"off"never tries UDP. TheYB_UDPvariable sets the default. UDP needs outbound UDP to the model server's port; see Security for how it's encrypted.
One robot's open session on one model. Use one per robot, from one thread at a time: its methods must not run at the same time.
| Member | What it does |
|---|---|
infer(observation, captured_ns=None) |
Sends one observation (camera images, arm state and task) and waits for the answer. Returns a dict: actions (the next moves, one row per step, as a NumPy array), client_action_age_ms (how old the answer was on arrival), sequence and epoch (which request and which episode it answers), model_revision (which model version answered) and timing. Raises InferenceError if the request fails. |
submit(observation, captured_ns=None) |
Sends one observation without waiting, and returns its sequence number. Collect the answer with result. Send the next observation while the robot still executes the current actions, so the model's time is hidden behind the motion. The model server keeps only the newest waiting observation per session, so submitting again before the last one started replaces it (superseded). |
result(sequence, timeout=None) |
Waits for the answer to a submitted observation, until its deadline, and returns the same dict as infer. With timeout (seconds) it waits less and raises TimeoutError, leaving the request pending, so your loop can do other work and ask again. |
capture_time_ns() |
The SDK's clock, in nanoseconds. Read it when the cameras take the pictures, and pass it as captured_ns. |
network() |
How the session reaches the model server: path ("udp" or "websocket"), udp_round_trip_ms (measured as the session opened), note (why UDP isn't in use, if it isn't), answers (how many answers arrived first by each path) and udp (packet counts, including resends). |
reset(instruction=None) |
Starts a new episode (a new attempt at the task), optionally with a new task. Answers meant for the old episode are dropped. |
status() |
The platform's view of the session: state, completed requests, cost so far, and the model server's counts. |
renew_credentials() |
Gets a fresh session pass now. You rarely need it: the SDK renews the pass once half its life has passed, and through a short platform outage it keeps trying until the pass is about to expire. |
report() |
The model server's report on the session: counts and timings. |
flush_telemetry() |
Sends up to 100 recent timings measured on your side, to help find problems. It never changes what you pay. |
close() |
Ends the session and settles its cost. Safe to call twice. Leaving a with block calls it. |
closure |
Set by close(): completed (requests answered), charged_usd, settlement_source (see Billing), close_confirmed and close_reason. close_reason is customer_closed when your code closed the session, or worker: plus the model server's reason when it had already ended the session, such as worker:too_many_expired. It's None if the platform couldn't be reached; the session then ends on its own, and nothing extra is charged. If the model server couldn't be reached, closure shows the state closing and charged_usd 0 until the platform settles it. |
assignment |
The platform's record of the session: model, version, the model server it runs on (generation), price, spending cap and request limit. |
id, model_id, model_revision, generation |
Which session this is, which model version it runs, and which model server it's on. |
Errors
Every error has a code, such as capacity. Errors says what each code
means and how to fix it.
| Class | Raised when | Extra members |
|---|---|---|
PlatformError |
a call to the platform fails. Codes not listed below, such as terms_required or email_unverified, raise a plain PlatformError. |
status (the HTTP status), context (extra details), retryable |
CapacityError |
no_capacity, capacity, session_limit, account_session_limit or rate_limit (a kind of PlatformError) |
the same |
AuthorizationError |
unauthorized, forbidden, not_found or rate_changed |
the same |
SpendError |
insufficient_credit, spend_too_small or spend_limit |
the same |
InferenceError |
infer or reset fails: for example a late answer (expired), a closed session, or a dropped connection (connection_lost). |
none |
Every error also has code, message, fix, retry and docs_url. retry says what
to do next: retry, fix, next_request, new_session or none. The table on
Errors explains each. Printing an error shows all of these.
For example, to retry when the model is full, waiting longer each time:
import time
from robot_inference_client import CapacityError
def open_session(client, instruction, attempts=5):
for attempt in range(attempts):
try:
return client.session("pi05-libero", instruction, max_spend_usd="5")
except CapacityError as error:
if not error.retryable or attempt == attempts - 1:
raise
time.sleep(2**attempt) # nothing is charged while you wait
Environment variables
| Variable | Meaning |
|---|---|
YB_URL |
Another copy of the platform to use instead of the hosted service. You don't need it otherwise. |
YB_API_KEY |
Your key, instead of the one saved by yb setup. For servers and CI. |
YB_PROJECT_ID |
Which project to use, when one key can see more than one. |
YB_ALLOW_INSECURE_HTTP |
Set to 1 to allow plain http to another machine on a private network you control. Without it, the SDK sends keys only over https, or to this machine. |
YB_UDP |
The default for the udp session option: auto, on or off. |
Robot safety
The SDK checks that each answer belongs to the request you sent, and that it's fresh when it arrives. It can't know whether the moves still fit the world when your robot makes them. Your robot's controller must check that sensor data is fresh, that each move still makes sense, and when to stop.