Every plan is a monthly inference allowance and a ceiling on concurrent GPU instances. Tokens are priced per model at the rates on the catalog; GPU-hours are passed through at the provider's rate with no markup. Exceeding the monthly allowance returns a typed 429 budget.exceeded, never a surprise invoice.
Every inference call, at the per-model rate the ledger recorded for it. GPU-hours for dedicated pods are metered separately and passed through.
Yes. A BYOK credential routes a model through your own account; those tokens land on your invoice and Nozzle charges only the platform fee.
Free carries no monthly credits, so inference is refused until you pick a plan. Sandbox keys still read everything.