Section 2

2. The four fulfillment modes

Two independent axes, not four separate features.

Whose credential pays the upstream — Nozzle's, yours, or nobody's (because we're running the weights ourselves and paying for GPU hours instead of tokens).

Who gets the capacity — shared across all customers, or dedicated to one tenant.

CredentialCapacityYou payStatus
A Platform passthroughNozzle'ssharedper tokenlive
B Shared self-hostednone (our pod)sharedper tokenlive
C Bring your own keyyourssharedplatform feelive
D Dedicated self-hostednone (your pod)dedicatedper GPU-hourlive

Subscription-versus-per-use is a pricing shape that can sit on any of these, not a fifth mode. Dedicated capacity is where a monthly floor makes sense, because you're holding a GPU whether you call it or not.

Mode A is the foundation: one key, every closed model, no accounts to manage. Mode B is where open weights become margin — Nozzle runs one pod, everyone calls it. Mode D is when you want a machine nobody else touches. Mode C is the override for when the tokens must land on your invoice.

A tenant-scoped binding always beats a global one, so registering your own model under a name the platform already serves silently overrides it — for you only.