Model pricing
Base models are priced per training token and per served token. Input and output tokens cost the same, so there is no prefill and sampling matrix to decode. All prices in USD per million tokens.
Price against capability
Every model in the catalogue, plotted by what it costs against what it returns. The dashed line joins the models nothing beats on both counts.
| Model | Prices |
|---|---|
| nvidia/Nemotron-3.5-LightningMoE · 12B params · 1.5B active | $0.22 train$0.15 infer128K |
| openai/gpt-oss-20bMoE · 21B params · 3.6B active | $0.36 train$0.25 infer128K |
| nvidia/Nemotron-3-Nano-30B-A3BMoE · 30B params · 3B active | $0.40 train$0.27 infer64K |
| Qwen/Qwen3-8BDense · 8B params | $0.40 train$0.30 infer32K |
| Qwen/Qwen3.8-27B-A3BMoE · 27B params · 3B active | $0.44 train$0.31 infer256K |
| Qwen/Qwen3.6-35B-A3BMoE · 35B params · 3B active | $0.48 train$0.34 infer256K |
| zai-org/GLM-5.3-FlashMoE · 106B params · 12B active | $0.60 train$0.40 infer1M |
| openai/gpt-oss-120bMoE · 117B params · 5.1B active | $0.68 train$0.45 infer128K |
| MiniMaxAI/MiniMax-M3MoE · 456B params · 46B active | $0.72 train$0.48 infer1M |
| mistralai/Mistral-Medium-3.5Dense · 41B params | $1.05 train$0.70 infer256K |
| deepseek-ai/DeepSeek-V4-Flash-0731MoE · 284B params · 13B active | $1.70 train$1.10 infer1M |
| deepseek-ai/DeepSeek-V4-FlashMoE · 284B params · 13B activePreview, superseded by 0731 · serve only · withdrawn 31 Jan 2027 | Not trainable$1.10 infer1M |
| meta-llama/Llama-3.3-70BDense · 70B params | $2.75 train$1.90 infer128K |
| deepseek-ai/DeepSeek-V3.2MoE · 685B params · 37B active | $3.20 train$2.10 infer128K |
| deepseek-ai/DeepSeek-V3.1MoE · 671B params · 37B activeWithdrawn 31 Jan 2027 | $3.40 train$2.30 infer128K |
| zai-org/GLM-5.2MoE · 753B params · 40B active | $3.95 train$2.55 infer1M |
| moonshotai/Kimi-K2.6MoE · 1T params · 32B activeWithdrawn 31 Jan 2027 | $4.40 train$2.75 infer128K |
| zai-org/GLM-5.3MoE · 780B params · 40B active | $4.60 train$2.95 infer1M |
| nvidia/Nemotron-3-Ultra-550B-A55BMoE · 550B params · 55B active | $5.00 train$3.10 infer64K |
| deepseek-ai/DeepSeek-V4-ProMoE · 1.6T params · 49B active | $5.60 train$3.50 infer1M |
| moonshotai/Kimi-K3MoE · 2.8T params · 104B active | Not trainable$6.50 infer1M |
| Model | Prices |
|---|---|
| google/gemma-4-E4BDense · 8B params · 4.5B effective | $0.35 train$0.24 infer128K |
| Qwen/Qwen3-VL-30B-A3B-InstructMoE · 30B params · 3B active | $0.45 train$0.32 infer128K |
| Qwen/Qwen3-VL-8B-InstructDense · 8B params | $0.48 train$0.35 infer128K |
| Qwen/Qwen3.5-9BDense · 9B params | $1.30 train$0.90 infer64K |
| google/gemma-4-31BDense · 31B params | $1.45 train$0.95 infer256K |
| Qwen/Qwen3-VL-235B-A22B-InstructMoE · 235B params · 22B active | $2.40 train$1.60 infer128K |
| Qwen/Qwen3.5-397B-A17BMoE · 397B params · 17B active | $6.00 train$3.75 infer64K |
| Model | Prices |
|---|---|
| google/embeddinggemma-300mBi-encoder · 300M params · 768 dims | $0.03 train$0.02 infer2K |
| BAAI/bge-reranker-v2-m3Cross-encoder · 568M params · 100+ languages | $0.04 train$0.03 infer8K |
| BAAI/bge-en-iclBi-encoder · 7B params · 4096 dims | $0.08 train$0.05 infer32K |
| Qwen/Qwen3-Embedding-8BBi-encoder · 8B params · 4096 dims | $0.08 train$0.05 infer32K |
| nomic-ai/nomic-embed-codeBi-encoder · 7B params · 3584 dims | $0.08 train$0.05 infer32K |
| Qwen/Qwen3-Reranker-8BCross-encoder · 8B params | $0.09 train$0.06 infer32K |
| Model | Prices |
|---|---|
| black-forest-labs/FLUX.2-kleinFlow transformer · 4B params | $0.50 train$2.50 inferUp to 2MP |
| stabilityai/stable-diffusion-3.5-largeMMDiT · 8B params | $0.90 train$4.50 inferUp to 2MP |
| Qwen/Qwen-ImageMMDiT · 20B params | $1.30 train$7.50 inferUp to 4MP |
| black-forest-labs/FLUX.2-devFlow transformer · 32B params | $1.40 train$8.00 inferUp to 4MP |
| Model | Prices |
|---|---|
| tencent/HunyuanVideo-1.5DiT · 8.3B params | $1.40 train$5.00 infer720p · 10s |
| Lightricks/LTX-2DiT · native audio + video | $1.60 train$5.50 inferUp to 4K · 10s |
| Wan-AI/Wan2.2-T2V-A14BMoE DiT · 27B params · 14B active | $1.80 train$6.50 infer720p · 5s |
| Model | Prices |
|---|---|
| openai/whisper-large-v3-turboASR · 809M params | $0.28 train$0.18 infer30s windows |
| nvidia/parakeet-tdt-1.1bASR · 1.1B params | $0.30 train$0.20 inferStreaming |
| nvidia/diar_sortformer_4spk-v1Diarisation · 115M params | $0.35 train$0.24 inferUp to 4 speakers |
| nvidia/Nemotron-3.5-ASR-MultilingualASR · 1.1B params · 25 languages | $0.45 train$0.30 inferStreaming |
| Qwen/Qwen3-TTS-1.7BTTS · 1.7B params · streaming | $0.50 train$2.50 infer10 languages |
| openai/whisper-large-v3ASR · 1.5B params | $0.60 train$0.40 infer30s windows |
| canopylabs/orpheus-3b-0.1-ftTTS · 3B params · streaming | $0.70 train$3.50 infer8 voices |
| mistralai/Voxtral-4B-TTS-2603TTS · 4B params | $0.90 train$4.50 infer9 languages |
Prices are in USD per million tokens and apply to LoRA fine-tuning and serving through the platform. Image, video and audio models meter the tokens of the encoded media: a megapixel image is roughly 4K tokens, a five-second 720p clip roughly 70K, and a minute of audio roughly 3K. Storage is $0.10 per GB per month, metered by the GB-hour. Mixture-of-experts models are priced by their active parameters. A model marked not trainable is served from the base weights only; you cannot fine-tune it yet.
Dedicated hosting
Reserve a GPU by the second and run unlimited tokens on it, with guaranteed performance and full control. Autoscale to zero, so idle capacity costs nothing. Prices are per GPU per hour.
| Hardware | On-demand / GPU-hr / Reserved / GPU-hr |
|---|---|
| NVIDIA L40S48 GB · Entry · small models and adapters | $1.99 On-demand / GPU-hr$1.49 Reserved / GPU-hr |
| NVIDIA A10080 GB · Mid workhorse | $4.19 On-demand / GPU-hr$3.14 Reserved / GPU-hr |
| NVIDIA H10080 GB · Flagship compute | $5.49 On-demand / GPU-hr$4.12 Reserved / GPU-hr |
| NVIDIA H200141 GB · High memory | $5.99 On-demand / GPU-hr$4.49 Reserved / GPU-hr |
| NVIDIA B200180 GB · Blackwell flagship | $8.99 On-demand / GPU-hr$6.74 Reserved / GPU-hr |
Billed per second with no minimum runtime. Cold-start is not charged. Reserved rates apply to 1 to 6 month commitments and run 15 to 25% below on-demand; the figures shown are the deepest tier.
Talk to salesStorage
Datasets, checkpoints and exported weights. Metered by the GB-hour, so you only pay while you keep them.
$0.10
per GB / month
Pay as you go
Usage-based from the first token, with the whole API included.
- Managed fine-tuning jobs
- The low-level training loop
- Evaluations
- OpenAI + Anthropic compatible serving
- Weight export, any checkpoint, any time
Enterprise
For organisations running fine-tuning at scale or under specific compliance requirements.
- Volume pricing
- Tailored onboarding and migration support
- SLA-backed dedicated support
- PDPA-aligned data handling
- Custom fine-tuned models built with our team
Run Nucleus in your own environment
Private and on-prem deployment from $50,000 / year platform licence. Your GPU and cloud infrastructure are billed separately. Built for data-residency and air-gapped requirements.
Talk to salesFrequently asked questions
Can't find the answer to your question? Contact our team and we will get back to you.