Pricing

Every plan is a monthly inference allowance and a ceiling on concurrent GPU instances. Tokens are priced per model at the rates on the catalog; GPU-hours are passed through at the provider's rate with no markup. Exceeding the monthly allowance returns a typed 429 budget.exceeded, never a surprise invoice.

Free

$0
inference credits / month
  • 1 concurrent GPU instance
  • Models up to 7B params
  • Every closed model, one key
  • GPU time at cost
  • 1,000 requests / day
Start free

Hobby

$20
inference credits / month
  • 1 concurrent GPU instance
  • Models up to 13B params
  • Every closed model, one key
  • GPU time at cost
  • 5,000 requests / day
Choose Hobby

Pro

Popular
$200
inference credits / month
  • 5 concurrent GPU instances
  • Models up to 70B params
  • Every closed model, one key
  • GPU time at cost
  • 50,000 requests / day
Choose Pro

Team

$1,000
inference credits / month
  • 20 concurrent GPU instances
  • Models up to 200B params
  • Every closed model, one key
  • GPU time at cost
  • 500,000 requests / day
Choose Team

Enterprise

$10,000
inference credits / month
  • Unlimited concurrent GPU instances
  • Models up to any size
  • Every closed model, one key
  • GPU time at cost
  • No daily request cap
Talk to us

What counts against the allowance?

Every inference call, at the per-model rate the ledger recorded for it. GPU-hours for dedicated pods are metered separately and passed through.

Can I bring my own provider key?

Yes. A BYOK credential routes a model through your own account; those tokens land on your invoice and Nozzle charges only the platform fee.

What happens on the free plan?

Free carries no monthly credits, so inference is refused until you pick a plan. Sandbox keys still read everything.