Model pricing

Base models are priced per training token and per served token. Input and output tokens cost the same, so there is no prefill and sampling matrix to decode. All prices in USD per million tokens.

Price against capability

Every model in the catalogue, plotted by what it costs against what it returns. The dashed line joins the models nothing beats on both counts.

CompanyAlibabaNVIDIADeepSeekOpenAIZ AIGoogleKimiMistralMiniMaxMetaOtherreads more than one modalityunscored, shown in the lanewithdrawal scheduled
Language
ModelPrices
nvidia/Nemotron-3.5-LightningMoE · 12B params · 1.5B active$0.22 train$0.15 infer128K
openai/gpt-oss-20bMoE · 21B params · 3.6B active$0.36 train$0.25 infer128K
nvidia/Nemotron-3-Nano-30B-A3BMoE · 30B params · 3B active$0.40 train$0.27 infer64K
Qwen/Qwen3-8BDense · 8B params$0.40 train$0.30 infer32K
Qwen/Qwen3.8-27B-A3BMoE · 27B params · 3B active$0.44 train$0.31 infer256K
Qwen/Qwen3.6-35B-A3BMoE · 35B params · 3B active$0.48 train$0.34 infer256K
zai-org/GLM-5.3-FlashMoE · 106B params · 12B active$0.60 train$0.40 infer1M
openai/gpt-oss-120bMoE · 117B params · 5.1B active$0.68 train$0.45 infer128K
MiniMaxAI/MiniMax-M3MoE · 456B params · 46B active$0.72 train$0.48 infer1M
mistralai/Mistral-Medium-3.5Dense · 41B params$1.05 train$0.70 infer256K
deepseek-ai/DeepSeek-V4-Flash-0731MoE · 284B params · 13B active$1.70 train$1.10 infer1M
deepseek-ai/DeepSeek-V4-FlashMoE · 284B params · 13B activePreview, superseded by 0731 · serve only · withdrawn 31 Jan 2027Not trainable$1.10 infer1M
meta-llama/Llama-3.3-70BDense · 70B params$2.75 train$1.90 infer128K
deepseek-ai/DeepSeek-V3.2MoE · 685B params · 37B active$3.20 train$2.10 infer128K
deepseek-ai/DeepSeek-V3.1MoE · 671B params · 37B activeWithdrawn 31 Jan 2027$3.40 train$2.30 infer128K
zai-org/GLM-5.2MoE · 753B params · 40B active$3.95 train$2.55 infer1M
moonshotai/Kimi-K2.6MoE · 1T params · 32B activeWithdrawn 31 Jan 2027$4.40 train$2.75 infer128K
zai-org/GLM-5.3MoE · 780B params · 40B active$4.60 train$2.95 infer1M
nvidia/Nemotron-3-Ultra-550B-A55BMoE · 550B params · 55B active$5.00 train$3.10 infer64K
deepseek-ai/DeepSeek-V4-ProMoE · 1.6T params · 49B active$5.60 train$3.50 infer1M
moonshotai/Kimi-K3MoE · 2.8T params · 104B activeNot trainable$6.50 infer1M
Vision
ModelPrices
google/gemma-4-E4BDense · 8B params · 4.5B effective$0.35 train$0.24 infer128K
Qwen/Qwen3-VL-30B-A3B-InstructMoE · 30B params · 3B active$0.45 train$0.32 infer128K
Qwen/Qwen3-VL-8B-InstructDense · 8B params$0.48 train$0.35 infer128K
Qwen/Qwen3.5-9BDense · 9B params$1.30 train$0.90 infer64K
google/gemma-4-31BDense · 31B params$1.45 train$0.95 infer256K
Qwen/Qwen3-VL-235B-A22B-InstructMoE · 235B params · 22B active$2.40 train$1.60 infer128K
Qwen/Qwen3.5-397B-A17BMoE · 397B params · 17B active$6.00 train$3.75 infer64K
Embedding & reranking
ModelPrices
google/embeddinggemma-300mBi-encoder · 300M params · 768 dims$0.03 train$0.02 infer2K
BAAI/bge-reranker-v2-m3Cross-encoder · 568M params · 100+ languages$0.04 train$0.03 infer8K
BAAI/bge-en-iclBi-encoder · 7B params · 4096 dims$0.08 train$0.05 infer32K
Qwen/Qwen3-Embedding-8BBi-encoder · 8B params · 4096 dims$0.08 train$0.05 infer32K
nomic-ai/nomic-embed-codeBi-encoder · 7B params · 3584 dims$0.08 train$0.05 infer32K
Qwen/Qwen3-Reranker-8BCross-encoder · 8B params$0.09 train$0.06 infer32K
Image generation
ModelPrices
black-forest-labs/FLUX.2-kleinFlow transformer · 4B params$0.50 train$2.50 inferUp to 2MP
stabilityai/stable-diffusion-3.5-largeMMDiT · 8B params$0.90 train$4.50 inferUp to 2MP
Qwen/Qwen-ImageMMDiT · 20B params$1.30 train$7.50 inferUp to 4MP
black-forest-labs/FLUX.2-devFlow transformer · 32B params$1.40 train$8.00 inferUp to 4MP
Video generation
ModelPrices
tencent/HunyuanVideo-1.5DiT · 8.3B params$1.40 train$5.00 infer720p · 10s
Lightricks/LTX-2DiT · native audio + video$1.60 train$5.50 inferUp to 4K · 10s
Wan-AI/Wan2.2-T2V-A14BMoE DiT · 27B params · 14B active$1.80 train$6.50 infer720p · 5s
Audio
ModelPrices
openai/whisper-large-v3-turboASR · 809M params$0.28 train$0.18 infer30s windows
nvidia/parakeet-tdt-1.1bASR · 1.1B params$0.30 train$0.20 inferStreaming
nvidia/diar_sortformer_4spk-v1Diarisation · 115M params$0.35 train$0.24 inferUp to 4 speakers
nvidia/Nemotron-3.5-ASR-MultilingualASR · 1.1B params · 25 languages$0.45 train$0.30 inferStreaming
Qwen/Qwen3-TTS-1.7BTTS · 1.7B params · streaming$0.50 train$2.50 infer10 languages
openai/whisper-large-v3ASR · 1.5B params$0.60 train$0.40 infer30s windows
canopylabs/orpheus-3b-0.1-ftTTS · 3B params · streaming$0.70 train$3.50 infer8 voices
mistralai/Voxtral-4B-TTS-2603TTS · 4B params$0.90 train$4.50 infer9 languages

Prices are in USD per million tokens and apply to LoRA fine-tuning and serving through the platform. Image, video and audio models meter the tokens of the encoded media: a megapixel image is roughly 4K tokens, a five-second 720p clip roughly 70K, and a minute of audio roughly 3K. Storage is $0.10 per GB per month, metered by the GB-hour. Mixture-of-experts models are priced by their active parameters. A model marked not trainable is served from the base weights only; you cannot fine-tune it yet.

Dedicated hosting

Reserve a GPU by the second and run unlimited tokens on it, with guaranteed performance and full control. Autoscale to zero, so idle capacity costs nothing. Prices are per GPU per hour.

HardwareOn-demand / GPU-hr / Reserved / GPU-hr
NVIDIA L40S48 GB · Entry · small models and adapters$1.99 On-demand / GPU-hr$1.49 Reserved / GPU-hr
NVIDIA A10080 GB · Mid workhorse$4.19 On-demand / GPU-hr$3.14 Reserved / GPU-hr
NVIDIA H10080 GB · Flagship compute$5.49 On-demand / GPU-hr$4.12 Reserved / GPU-hr
NVIDIA H200141 GB · High memory$5.99 On-demand / GPU-hr$4.49 Reserved / GPU-hr
NVIDIA B200180 GB · Blackwell flagship$8.99 On-demand / GPU-hr$6.74 Reserved / GPU-hr

Billed per second with no minimum runtime. Cold-start is not charged. Reserved rates apply to 1 to 6 month commitments and run 15 to 25% below on-demand; the figures shown are the deepest tier.

Talk to sales

Storage

Datasets, checkpoints and exported weights. Metered by the GB-hour, so you only pay while you keep them.

$0.10

per GB / month

Pay as you go

Usage-based from the first token, with the whole API included.

  • Managed fine-tuning jobs
  • The low-level training loop
  • Evaluations
  • OpenAI + Anthropic compatible serving
  • Weight export, any checkpoint, any time
Get started

Enterprise

For organisations running fine-tuning at scale or under specific compliance requirements.

  • Volume pricing
  • Tailored onboarding and migration support
  • SLA-backed dedicated support
  • PDPA-aligned data handling
  • Custom fine-tuned models built with our team
Request a demo

Run Nucleus in your own environment

Private and on-prem deployment from $50,000 / year platform licence. Your GPU and cloud infrastructure are billed separately. Built for data-residency and air-gapped requirements.

Talk to sales

Frequently asked questions

Can't find the answer to your question? Contact our team and we will get back to you.

Your models, served as a managed API. At a fraction of the cost.