Your own GPU, billed by the second
Reserve an L40S, A100, H100, H200 or B200 and run unlimited tokens on it, with guaranteed performance and full control. Cold starts are free, and idle capacity costs nothing.
Billed by the second
There is no minimum runtime. The meter counts the seconds you hold the card and stops when you release it.
Unlimited tokens
Run as many tokens as the card can produce. The GPU-hour is the only meter on dedicated hosting.
Cold starts free
A cold start is never charged. Autoscale to zero, and idle capacity costs nothing.
Guaranteed performance
The card is yours while you hold it, with guaranteed performance and full control over what runs on it.
What a dedicated GPU gives you
Pick the card, hold it by the second, open a session through the API and watch the day's load.
The L40S carries small models and adapters. The A100 and H100 carry 80 GB, the H200 141 GB and the B200 180 GB, so the card you pick is the memory your model needs. Every price is per GPU per hour, and every hour is billed by the second.
Results from LLMs powered by Nucleus
daily fuel saved across the fleet
sweep of the tracking feed
from site reading to client report
enclosure sensors on the feed
Workloads that run all day
47 trucks rerouted while they were still moving
Every shipment two days past ETA, found before the customer called
Cable survey readings processed while the crew was still on site
At a glance
Hardware
- NVIDIA L40S, 48 GB
- NVIDIA A100, 80 GB
- NVIDIA H100, 80 GB
- NVIDIA H200, 141 GB
- NVIDIA B200, 180 GB
Billing
- Per second, no minimum runtime
- Cold start never charged
- Unlimited tokens on the card
- Autoscale to zero
- Prices per GPU per hour
Serving
- OpenAI-compatible route
- Anthropic-compatible route
- Samplers API: create, retrieve, close
- Base model or your nucleus:// weights
- The infer scope
Terms
- On-demand $1.99 to $8.99 / GPU-hr
- Reserved 15 to 25% below on-demand
- One to six month commitments
- Billed separately from tokens and storage
- Talk to sales for reserved
Dedicated GPUs start at $1.99 per GPU-hour
L40S to B200, billed by the second. Reserved rates apply to 1 to 6 month commitments and run 15 to 25% below on-demand; the figures shown are the deepest tier.
Dedicated GPUs
L40S to B200, billed by the second.
Pick your path
Private Endpoints
OpenAI and Anthropic compatible
Serverless
Pay per token, idle costs nothing
Reserved capacity
Guaranteed terms, one to six months
Compatible APIs
DocsOne key, two SDK families
Resources
Guide
Take finished weights to production by serving them on Nucleus or exporting them to run anywhere.
API reference
Open inference sessions against base models or your own fine-tuned weights, and generate from raw token IDs.
Customer story
Mandai Wildlife Group transforms animal care operations with Nucleus
Reserve a GPU by the second
L40S to B200, unlimited tokens, cold starts free.