Yellow and Black

Docs Start here

Quickstart

Yellow and Black runs AI models for robots on our GPUs. Today it serves vision-language-action (VLA) models, such as π0.5. It lets you:

  • Open a session; we prepare the selected model and connect you when it is ready
  • Run inference for a real robot or a simulation, with one of our models

Set up

  1. Create an account at https://yellowandblack.dev/platform, accept the service terms, and verify your email to receive the one-time $5 trial credit. Then create an API key in the API keys tab.

  2. Use a virtual environment for your project. If you already have one, activate it. Otherwise, create one in your project folder. On macOS or Linux:

    python3 -m venv .venv
    source .venv/bin/activate
    

    On Windows PowerShell:

    py -m venv .venv
    .\.venv\Scripts\Activate.ps1
    

    If PowerShell blocks activation, use .\.venv\Scripts\python.exe instead of python and .\.venv\Scripts\yb.exe instead of yb in the commands below. Activate the environment again when you open a new terminal.

    Install the SDK and add your key. Paste the key when asked, and answer y to save it.

    python -m pip install "robot-inference-client[keyring] @ https://yellowandblack.dev/sdk/robot_inference_client-0.5.3-py3-none-any.whl"
    yb setup --url "https://yellowandblack.dev"
    

    If a supported OS keyring is available, setup offers to save your key there. Otherwise, follow its instructions to set YB_API_KEY in your shell.

    Use this installation's address for later CLI commands. In bash/zsh:

    export YB_URL="https://yellowandblack.dev"
    

    In PowerShell:

    $env:YB_URL = "https://yellowandblack.dev"
    
  3. See which models are offered:

    yb models --ready
    

    An offered model may be idle. Opening a session starts it; listing models does not start one or reserve capacity.

Send your first request

The SDK includes a recorded LIBERO simulator observation. No separate download or source checkout is needed. Send it to an offered pi05-libero model:

from robot_inference_client import Client, example_observation

task, observation = example_observation()
with Client(base_url="https://yellowandblack.dev") as client, client.session(
    model="pi05-libero", instruction=task, max_spend_usd="1"
) as session:
    result = session.infer(observation)
    print(result["actions"].shape)  # (10, 7): the next 10 actions

Save it as first_call.py, and run it:

python first_call.py

Opening the session waits for the model to become ready. Add on_status=print to client.session(...) to display starting and ready when those states are reached. Startup and waiting are free; they are separate from inference time.

If the SDK's startup waiting limit expires, it raises model_start_timeout and requests cancellation. By default it uses the model's advertised waiting limit. You can set a shorter limit with connect_timeout; this does not make the model start faster. See Session options for checking whether cancellation was confirmed before opening another session.

Note: max_spend_usd="1" caps the session at $1 of available credit. This is replay input: its send time is not a camera capture time, and its output does not demonstrate task success. The recorded example uses the LIBERO schema, not DROID.

The same example is available from the terminal:

yb run --model pi05-libero --example --requests 1 --max-spend-usd 1

For your own recorded LIBERO observation, use --fixture observation.npz instead of --example. A failed inference preserves its JSON report and exits nonzero; temporary retryable failures use exit status 75.

Test your setup for free

transport-demo is a free sandbox. It isn't a model: it returns placeholder actions in the same format, so you can check your key, network and code without spending credit.

from robot_inference_client import Client
from robot_inference_client.cli import synthetic_observation

with Client(base_url="https://yellowandblack.dev") as client, client.session(
    model="transport-demo", instruction="test"
) as session:
    result = session.infer(synthetic_observation(0, "test"))
    print(result["actions"].shape)  # (10, 7): placeholder actions

Or check it from a terminal. This sends 10 placeholder observations, then prints how long each answer took and what the session cost:

yb run --model transport-demo

For a DROID robot, use yb run --model transport-demo-droid.

Run it in a control loop

Replace the placeholder images with data from your robot or simulator, and call infer once per step:

from robot_inference_client import Client, InferenceError

task = "put the bowl on the plate"
with Client(base_url="https://yellowandblack.dev") as client, client.session(
    model="pi05-libero",
    instruction=task,
    max_spend_usd="5",
    transport="owned",  # bounds the SDK's socket waits
) as session:
    while robot.running():
        captured = session.capture_time_ns()  # when the cameras take the pictures
        observation = {
            "observation/image": robot.main_camera(),  # uint8, (224, 224, 3)
            "observation/wrist_image": robot.wrist_camera(),  # uint8, (224, 224, 3)
            "observation/state": robot.state(),  # 8 numbers: arm position and gripper
            "prompt": task,
        }
        try:
            result = session.infer(observation, captured_ns=captured)
        except InferenceError as error:
            if error.retry == "next_request":
                continue  # this answer came too late or was replaced: send the next one
            raise
        robot.execute(result["actions"])  # the next moves

Each model expects its own data format; Robot data formats lists them.

Next steps

  • Python SDK: every class and method, and each session option
  • Billing: prices, credit and spending caps
  • Errors: what each error means, and how to fix it

Updated 2026-10-01 · View as markdown