Questions and answers
What is Yellow and Black?
A service that runs AI models for robots on our GPUs. Today it serves vision-language-action (VLA) models, such as π0.5 from Physical Intelligence. You call them from a real robot or a simulation, and you don't rent, set up or load a GPU.
What is a session?
One robot's connection to one model, from when you open it to when you close it. A session isn't a whole GPU: a few sessions share one model server and take turns. Each session's pass works only for that session.
What do I pay for?
Each request the model answers. Failed requests are free, and so are loading and idle time. You set a spending cap for each session. See Billing.
Do I need a GPU?
No. The model runs on our GPUs. Your robot's computer needs Python 3.8 or newer and an internet connection.
Which robots and models are supported?
yb models --ready lists the offered models and their current state. An on-demand
model may be idle; opening a session waits for startup before inference can begin.
Today that is π0.5, in two robot setups:
- LIBERO: a simulated Franka robot arm with two cameras. It sends 8 numbers for the arm's position and gripper, and gets back 7 numbers per move.
- DROID: a Franka robot arm steered by joint speeds. Its model has run on our GPUs, but hasn't been tested with a real robot yet.
To test your code for free, use transport-demo: a sandbox, not a model, that returns
placeholder actions. See Robot data formats.
How fast is it?
In private tests on October 1, 2026, L40S GPUs served two simultaneous clients per model. Median SDK response time across 500 shared requests per model was 169 ms for LIBERO and 172 ms for DROID. The caller was in a US-East datacenter; this does not measure the network between your robot and the service. Shared capacity and your connection affect response time.
Opening an idle model is a separate wait. In 15 cold-start trials per model, successful connections had medians of 12 seconds for LIBERO and 15 seconds for DROID. The slowest successful connections took 110 and 97 seconds, and one LIBERO trial timed out. These results do not establish a startup guarantee; capacity delays remain unresolved. Startup and waiting are free.
Pass on_status=print to client.session(...) to see connection progress. See
Session options for the waiting limit and
startup-error handling.
What if my network is slow or drops?
The model server drops an answer that would be too late, and doesn't charge for it. An
answer it sent in time that reached you late is still charged. Either way, your code gets
an expired error and sends the next observation. If the connection drops, the
SDK raises connection_lost and closes the session: open a new one. Your robot's controller decides what the robot does while
it waits.
Do you store my camera images?
Our gateway does not deliberately save camera images or robot state. For Modal-backed inference, Modal may retain function inputs and outputs for up to seven days, even when diagnostic capture is disabled. Account deletion does not immediately erase those provider copies. We keep session counts, timings and costs. See Security and data and Modal's retention policy.
Can AI coding agents use it?
Yes. Every docs page is also available as markdown, /llms.txt lists them all,
and every error has a code. Through the SDK and yb, each error also comes with a fix and
a link. See Guide for AI agents.
Can I train or fine-tune a model here?
Not yet. You can use the models that yb models --ready lists.
Is it available now?
This release is being prepared for public launch. Public signup and hosted model
availability have not been verified. On a configured private installation, the
free transport-demo sandbox lets you check your integration with placeholder
actions; it does not qualify a real model or robot task.