# Billing and credit

You pay per request the model answers. You pay from prepaid credit, and each session has a
spending cap that you set.

## Prices

Each model has a price per 1,000 requests. The **Models** tab shows it, and so does
`yb models --ready`, under `rate`.

- `transport-demo` is free. It's a sandbox for testing your setup, not a model: it returns
  placeholder actions.
- Using a model means you accept its price. There's nothing to click first.
- Check the published rate before opening a paid session. Open sessions keep their
  original rate. Supplying `rate_version` rejects a new session if its rate has changed.

## How a session is billed

1. **Start with trial credit.** The launch offer gives each eligible new user **$5**
   in promotional credit once, automatically after email verification and acceptance
   of the current service terms. There is no total giveaway cap. Deleting an account
   and registering again with the same email does not grant another trial.
   To purchase more credit, open **Usage & billing**; the configured purchase range is displayed
   there (default $25–$100). Promotional credit is spent first. Unused promotional
   and purchased credits are nonrefundable and are forfeited when you confirm account
   deletion.
2. **Open a session with a cap.** `max_spend_usd` is the most the session can cost. That
   amount is reserved from your credit while the session is open.
3. **The cap limits requests.** A session can make at most its cap divided by the price of
   one request, rounded down. At that point it stops with `spend_exhausted`. Answers the
   model server drops as late aren't charged, but they still count toward this limit,
   because the model ran for them.
4. **Close the session.** You pay the requests the model answered, times the price, never
   more than the cap. The unused reservation is released when settlement completes.
   An unreachable worker can leave the session closing while cleanup retries. Each
   session is settled exactly once.

The free sandbox needs no credit, and has no request limit.

An on-demand model may show **Starting** while capacity is prepared. Your spending
cap is held during this wait, but no inference is charged. Cancelling or timing out
before admission releases the hold. Other customers using the shared model keep
their connections. After the last connection ends, we may shut down its GPU.

## Example

For example, at a made-up price of $2.00 per 1,000 requests:

- `max_spend_usd="5"` reserves $5.00 and allows up to 2,500 requests.
- A robot that sends ten requests a second for three minutes sends 1,800 requests. If the
  model answers all of them, the session costs $3.60.
- When the session closes, $3.60 is charged and $1.40 comes back to your credit.

## What's free

- Opening a session, loading a model, starting it up, and idle time.
- Requests refused because of their format, the key, a full model, or a rate limit.
- Answers the model server drops because they would be too late (`expired`), because
  newer data replaced them (`superseded`), or because the model failed (`worker_failed`).
  An answer the server sent in time, but that reached your code late, is still charged.
- Unfinished work when a model server is replaced (`replica_retired`).

## Refunds, disputes and lost servers

- A refunded or disputed card payment removes that credit at once. If the remaining
  credit no longer covers what open sessions reserved, they close with `credit_reversed`,
  and are settled for what they completed.
- A released dispute hold restores only the corresponding held credit. Settled refunds
  remain removed.
- If the machine running a model server is lost, its open sessions are charged the last
  count the server reported, never more than their caps. With no report, they're charged
  nothing.

## See your usage

- `yb usage`, or `client.usage()`, shows your balance, reserved credit and ledger.
- In the **Usage & billing** tab, **Export CSV** downloads the ledger.
  The console checks that the complete generated file arrived before offering
  the download. The report is generated live, so concurrent changes may not be
  included. Retry a reported verification failure; contact support if the report
  exceeds the 8 MiB download limit.
- **Purchases** shows pending, completed, expired, and review states. Leaving the
  checkout page does not confirm or cancel a payment. Resume its link or cancel the
  unfinished purchase from this history.
- Checkout retries reuse the same purchase identifier. If a response is lost, check
  the existing purchase before starting a new one. A payment ID marked for review
  lets support reconcile an uncertain provider outcome.
- Account deletion disables access first and continues cleanup in the background.
  Keep the deletion receipt to check progress. Existing unused credits are forfeited
  after settlement; a payment that completes during deletion is held for separate
  reconciliation and never restores account access.
- A closed session shows `completed` (requests answered), `charged_usd` and
  `settlement_source`, which says how the cost was settled:
  - `worker`: by the model server's own count
  - `snapshot`: by the model server's final count, saved when the server was replaced
  - `last_report`: the machine was lost, so the server's last reported count was used
  - `unavailable`: nothing was known, so nothing was charged

> **Note:** A charged request means our server sent the answer. It doesn't prove that your
> robot ran it, or that it still fit when it did. Timings your code reports help find
> problems, but the model server's count settles the bill.
