SERVE · DEDICATED GPUS

Your own GPU, billed by the second

Reserve an L40S, A100, H100, H200 or B200 and run unlimited tokens on it, with guaranteed performance and full control. Cold starts are free, and idle capacity costs nothing.

GPUsbilled by the second
4 types
L40S48 GB
$1.99 / GPU-hour
A10080 GB
$2.89 / GPU-hour
H10080 GB
$4.49 / GPU-hour
B200192 GB
$7.90 / GPU-hour
ep_7Kd1pH100 · one replica
serving
Utilisation
78%
142 tok/sthroughput
$0.00cold start

Billed by the second

There is no minimum runtime. The meter counts the seconds you hold the card and stops when you release it.

Unlimited tokens

Run as many tokens as the card can produce. The GPU-hour is the only meter on dedicated hosting.

Cold starts free

A cold start is never charged. Autoscale to zero, and idle capacity costs nothing.

Guaranteed performance

The card is yours while you hold it, with guaranteed performance and full control over what runs on it.

DEDICATED GPUS

What a dedicated GPU gives you

Pick the card, hold it by the second, open a session through the API and watch the day's load.

The L40S carries small models and adapters. The A100 and H100 carry 80 GB, the H200 141 GB and the B200 180 GB, so the card you pick is the memory your model needs. Every price is per GPU per hour, and every hour is billed by the second.

ServerlessIdle time costs nothing.
from $0.15/ 1M tokens
DedicatedA GPU by the second, L40S to B200.
from $1.99/ GPU-hour
ReservedGuaranteed terms, one to six months.
from $1.49/ GPU-hour
See the full ladder
Results

Results from LLMs powered by Nucleus

$2.4K

daily fuel saved across the fleet

Daily

sweep of the tracking feed

Same day

from site reading to client report

Live

enclosure sensors on the feed

USE CASES

Workloads that run all day

Fleet operations

47 trucks rerouted while they were still moving

$2.4Kdaily fuel saved across the fleet
23delays resolved without a call
Logistics

Every shipment two days past ETA, found before the customer called

Dailysweep of the tracking feed
2 dayspast ETA is the escalation line
Field engineering

Cable survey readings processed while the crew was still on site

Same dayfrom site reading to client report
9systems on one thread

At a glance

Hardware

  • NVIDIA L40S, 48 GB
  • NVIDIA A100, 80 GB
  • NVIDIA H100, 80 GB
  • NVIDIA H200, 141 GB
  • NVIDIA B200, 180 GB

Billing

  • Per second, no minimum runtime
  • Cold start never charged
  • Unlimited tokens on the card
  • Autoscale to zero
  • Prices per GPU per hour

Serving

  • OpenAI-compatible route
  • Anthropic-compatible route
  • Samplers API: create, retrieve, close
  • Base model or your nucleus:// weights
  • The infer scope

Terms

  • On-demand $1.99 to $8.99 / GPU-hr
  • Reserved 15 to 25% below on-demand
  • One to six month commitments
  • Billed separately from tokens and storage
  • Talk to sales for reserved
PRICING

Dedicated GPUs start at $1.99 per GPU-hour

L40S to B200, billed by the second. Reserved rates apply to 1 to 6 month commitments and run 15 to 25% below on-demand; the figures shown are the deepest tier.

Dedicated GPUs

$1.99/ GPU-hour

L40S to B200, billed by the second.

Pick your path

Pick your path

Private Endpoints

OpenAI and Anthropic compatible

Explore Private Endpoints

Serverless

Pay per token, idle costs nothing

Explore Serverless

Reserved capacity

Guaranteed terms, one to six months

Explore Reserved capacity

Compatible APIs

Docs

One key, two SDK families

Explore Compatible APIs

Reserve a GPU by the second

L40S to B200, unlimited tokens, cold starts free.