Reserve the capacity, keep the saving
Commit to a GPU for one to six months and pay 15 to 25% below on-demand for it. The capacity is guaranteed for the term, and the card runs unlimited tokens with guaranteed performance and full control.
One to six month terms
Reserve for a single month or for six. The rate is fixed for the term you choose.
15 to 25% below on-demand
A one-month term runs 15% under the on-demand rate and a six-month term runs 25% under it, on every card in the ladder.
Capacity guaranteed
The GPUs you reserve are held for you for the whole term.
The same card, unlimited tokens
A reserved GPU is a dedicated GPU. Run unlimited tokens on it, with guaranteed performance and full control.
What a reservation gives you
Pick a term, compare it with on-demand, read what the guarantee covers and send the reservation.
Six rungs, from one month to six
A one-month term runs 15% below on-demand and a six-month term runs 25% below it. The terms between are priced by their length. The rate you agree holds for the whole term, on the card you chose.
Results from LLMs powered by Nucleus
downtime prevented this month · Manufacturing
average time to resolution · IT operations
daily fuel saved across the fleet · Fleet operations
sweep of the tracking feed · Logistics
Workloads that run every day
Line 3's next failure, booked into the maintenance window
Every incident past its four-hour SLA, escalated by name
47 trucks rerouted while they were still moving
At a glance
Terms
- One to six months
- 15% below on-demand at one month
- 25% below on-demand at six months
- Capacity guaranteed for the term
- Talk to sales to reserve
Hardware
- NVIDIA L40S, 48 GB
- NVIDIA A100, 80 GB
- NVIDIA H100, 80 GB
- NVIDIA H200, 141 GB
- NVIDIA B200, 180 GB
Billing
- Prices per GPU per hour
- Unlimited tokens on the card
- Guaranteed performance and full control
- Billed separately from tokens and storage
- Cold start never charged
Serving
- OpenAI-compatible route
- Anthropic-compatible route
- Base model or your nucleus:// weights
- Samplers API: create, retrieve, close
- The infer scope
Reserved GPUs start at $1.49 per GPU-hour
Reserved rates apply to 1 to 6 month commitments and run 15 to 25% below on-demand; the figures shown are the deepest tier.
Reserved GPUs
L40S to B200, one to six month terms.
Pick your path
The other ways to serve a private endpoint.
Private Endpoints
OpenAI and Anthropic compatible
Serverless
Pay per token, idle costs nothing
Dedicated GPUs
L40S to B200, by the second
Compatible APIs
DocsOne key, two SDK families
Resources
Reserve the capacity you need
One to six month terms, 15 to 25% below on-demand, capacity guaranteed.