SERVE · RESERVED CAPACITY

Reserve the capacity, keep the saving

Commit to a GPU for one to six months and pay 15 to 25% below on-demand for it. The capacity is guaranteed for the term, and the card runs unlimited tokens with guaranteed performance and full control.

Reservationproj_acme · H100 × 4
active
$1.49per GPU-hour, held
25%under the by-the-second rate
Term
1 to 6 months
Capacity
guaranteed for the term
Overage
billed at the standard rate

One to six month terms

Reserve for a single month or for six. The rate is fixed for the term you choose.

15 to 25% below on-demand

A one-month term runs 15% under the on-demand rate and a six-month term runs 25% under it, on every card in the ladder.

Capacity guaranteed

The GPUs you reserve are held for you for the whole term.

The same card, unlimited tokens

A reserved GPU is a dedicated GPU. Run unlimited tokens on it, with guaranteed performance and full control.

RESERVED CAPACITY

What a reservation gives you

Pick a term, compare it with on-demand, read what the guarantee covers and send the reservation.

01 · TERMS

Six rungs, from one month to six

A one-month term runs 15% below on-demand and a six-month term runs 25% below it. The terms between are priced by their length. The rate you agree holds for the whole term, on the card you chose.

ServerlessIdle time costs nothing.
from $0.15/ 1M tokens
DedicatedA GPU by the second, L40S to B200.
from $1.99/ GPU-hour
ReservedGuaranteed terms, one to six months.
from $1.49/ GPU-hour
See the hosting ladder
Results

Results from LLMs powered by Nucleus

$177K

downtime prevented this month · Manufacturing

6.2 hrs

average time to resolution · IT operations

$2.4K

daily fuel saved across the fleet · Fleet operations

Daily

sweep of the tracking feed · Logistics

USE CASES

Workloads that run every day

Manufacturing

Line 3's next failure, booked into the maintenance window

$177Kdowntime prevented this month
18.3 daysadvance notice on a failure
See the use case
IT operations

Every incident past its four-hour SLA, escalated by name

8incidents past their four-hour SLA
6.2 hrsaverage time to resolution
See the use case
Fleet operations

47 trucks rerouted while they were still moving

$2.4Kdaily fuel saved across the fleet
23delays resolved without a call
See the use case

At a glance

Terms

  • One to six months
  • 15% below on-demand at one month
  • 25% below on-demand at six months
  • Capacity guaranteed for the term
  • Talk to sales to reserve

Hardware

  • NVIDIA L40S, 48 GB
  • NVIDIA A100, 80 GB
  • NVIDIA H100, 80 GB
  • NVIDIA H200, 141 GB
  • NVIDIA B200, 180 GB

Billing

  • Prices per GPU per hour
  • Unlimited tokens on the card
  • Guaranteed performance and full control
  • Billed separately from tokens and storage
  • Cold start never charged

Serving

  • OpenAI-compatible route
  • Anthropic-compatible route
  • Base model or your nucleus:// weights
  • Samplers API: create, retrieve, close
  • The infer scope
PRICING

Reserved GPUs start at $1.49 per GPU-hour

Reserved rates apply to 1 to 6 month commitments and run 15 to 25% below on-demand; the figures shown are the deepest tier.

Reserved GPUs

$1.49/ GPU-hour

L40S to B200, one to six month terms.

Pick your path

Pick your path

The other ways to serve a private endpoint.

Private Endpoints

OpenAI and Anthropic compatible

Explore Private Endpoints

Serverless

Pay per token, idle costs nothing

Explore Serverless

Dedicated GPUs

L40S to B200, by the second

Explore Dedicated GPUs

Compatible APIs

Docs

One key, two SDK families

Explore Compatible APIs
RESERVED CAPACITY

Reserve the capacity you need

One to six month terms, 15 to 25% below on-demand, capacity guaranteed.