2. The four fulfillment modes
Two independent axes, not four separate features.
Whose credential pays the upstream — Nozzle's, yours, or nobody's (because we're running the weights ourselves and paying for GPU hours instead of tokens).
Who gets the capacity — shared across all customers, or dedicated to one tenant.
| Credential | Capacity | You pay | Status | |
|---|---|---|---|---|
| A Platform passthrough | Nozzle's | shared | per token | live |
| B Shared self-hosted | none (our pod) | shared | per token | live |
| C Bring your own key | yours | shared | platform fee | live |
| D Dedicated self-hosted | none (your pod) | dedicated | per GPU-hour | live |
Subscription-versus-per-use is a pricing shape that can sit on any of these, not a fifth mode. Dedicated capacity is where a monthly floor makes sense, because you're holding a GPU whether you call it or not.
Mode A is the foundation: one key, every closed model, no accounts to manage. Mode B is where open weights become margin — Nozzle runs one pod, everyone calls it. Mode D is when you want a machine nobody else touches. Mode C is the override for when the tokens must land on your invoice.
A tenant-scoped binding always beats a global one, so registering your own model under a name the platform already serves silently overrides it — for you only.