Model library

Every base model on the platform, with the context window it takes, the highest LoRA rank it accepts, and what it costs to fine-tune and to serve. Prices are USD per million tokens.

33 models

Language

16
openai

GPT-OSS 20B

openai/gpt-oss-20b

MoE · 21B params · 3.6B active

Context
128K
LoRA rank
64
Fine-tune
$0.36
Serve
$0.25
nvidia

Nemotron 3 Nano 30B

nvidia/Nemotron-3-Nano-30B-A3B

MoE · 30B params · 3B active

Context
64K
LoRA rank
64
Fine-tune
$0.40
Serve
$0.27
Qwen

Qwen3 8B

Qwen/Qwen3-8B

Dense · 8B params

Context
32K
LoRA rank
64
Fine-tune
$0.40
Serve
$0.30
Qwen

Qwen3.6 35B

Qwen/Qwen3.6-35B-A3B

MoE · 35B params · 3B active

Context
256K
LoRA rank
64
Fine-tune
$0.48
Serve
$0.34
openai

GPT-OSS 120B

openai/gpt-oss-120b

MoE · 117B params · 5.1B active

Context
128K
LoRA rank
32
Fine-tune
$0.68
Serve
$0.45
nvidia

Nemotron 3 Super 120B

nvidia/Nemotron-3-Super-120B-A12B

MoE · 120B params · 12B active

Context
64K
LoRA rank
32
Fine-tune
$1.15
Serve
$0.72
Qwen

Qwen3 32B

Qwen/Qwen3-32B

Dense · 32B params

Context
32K
LoRA rank
32
Fine-tune
$1.35
Serve
$0.90
deepseek-ai

DeepSeek V4 Flash 0731

deepseek-ai/DeepSeek-V4-Flash-0731

MoE · 284B params · 13B active

Context
1M
LoRA rank
32
Fine-tune
$1.70
Serve
$1.10
deepseek-ai

DeepSeek V4 Flash

deepseek-ai/DeepSeek-V4-Flash

MoE · 284B params · 13B active

Preview, superseded by 0731 · withdrawn 31 Jan 2027

Context
1M
LoRA rank
32
Fine-tune
$1.70
Serve
$1.10
meta-llama

Llama 3.3 70B

meta-llama/Llama-3.3-70B

Dense · 70B params

Context
128K
LoRA rank
32
Fine-tune
$2.75
Serve
$1.90
deepseek-ai

DeepSeek V3.1

deepseek-ai/DeepSeek-V3.1

MoE · 671B params · 37B active

Withdrawn 31 Jan 2027

Context
128K
LoRA rank
16
Fine-tune
$3.40
Serve
$2.30
zai-org

GLM 5.2

zai-org/GLM-5.2

MoE · 753B params · 40B active

Context
1M
LoRA rank
16
Fine-tune
$3.95
Serve
$2.55
moonshotai

Kimi K2.6

moonshotai/Kimi-K2.6

MoE · 1T params · 32B active

Withdrawn 31 Jan 2027

Context
128K
LoRA rank
16
Fine-tune
$4.40
Serve
$2.75
nvidia

Nemotron 3 Ultra 550B

nvidia/Nemotron-3-Ultra-550B-A55B

MoE · 550B params · 55B active

Context
64K
LoRA rank
16
Fine-tune
$5.00
Serve
$3.10
deepseek-ai

DeepSeek V4 Pro

deepseek-ai/DeepSeek-V4-Pro

MoE · 1.6T params · 49B active

Context
1M
LoRA rank
16
Fine-tune
$5.60
Serve
$3.50
moonshotai
Serve only

Kimi K3

moonshotai/Kimi-K3

MoE · 2.8T params · 104B active

Context
1M
LoRA rank
n/a
Fine-tune
n/a
Serve
$6.50

Vision

7
google

Gemma 4 E4B

google/gemma-4-E4B

Dense · 8B params · 4.5B effective

Context
128K
LoRA rank
64
Fine-tune
$0.35
Serve
$0.24
Qwen

Qwen3-VL 30B Instruct

Qwen/Qwen3-VL-30B-A3B-Instruct

MoE · 30B params · 3B active

Context
128K
LoRA rank
64
Fine-tune
$0.45
Serve
$0.32
Qwen

Qwen3-VL 8B Instruct

Qwen/Qwen3-VL-8B-Instruct

Dense · 8B params

Context
128K
LoRA rank
64
Fine-tune
$0.48
Serve
$0.35
Qwen

Qwen3.5 9B

Qwen/Qwen3.5-9B

Dense · 9B params

Context
64K
LoRA rank
64
Fine-tune
$1.30
Serve
$0.90
google

Gemma 4 31B

google/gemma-4-31B

Dense · 31B params

Context
256K
LoRA rank
32
Fine-tune
$1.45
Serve
$0.95
Qwen

Qwen3-VL 235B Instruct

Qwen/Qwen3-VL-235B-A22B-Instruct

MoE · 235B params · 22B active

Context
128K
LoRA rank
32
Fine-tune
$2.40
Serve
$1.60
Qwen

Qwen3.5 397B

Qwen/Qwen3.5-397B-A17B

MoE · 397B params · 17B active

Context
64K
LoRA rank
32
Fine-tune
$6.00
Serve
$3.75

Image generation

4
black-forest-labs

FLUX.2 klein

black-forest-labs/FLUX.2-klein

Flow transformer · 4B params

Context
Up to 2MP
LoRA rank
64
Fine-tune
$0.50
Serve
$2.50
stabilityai

Stable Diffusion 3.5 Large

stabilityai/stable-diffusion-3.5-large

MMDiT · 8B params

Context
Up to 2MP
LoRA rank
64
Fine-tune
$0.90
Serve
$4.50
Qwen

Qwen-Image

Qwen/Qwen-Image

MMDiT · 20B params

Context
Up to 4MP
LoRA rank
32
Fine-tune
$1.30
Serve
$7.50
black-forest-labs

FLUX.2 dev

black-forest-labs/FLUX.2-dev

Flow transformer · 32B params

Context
Up to 4MP
LoRA rank
32
Fine-tune
$1.40
Serve
$8.00

Video generation

3
tencent

HunyuanVideo 1.5

tencent/HunyuanVideo-1.5

DiT · 8.3B params

Context
720p · 10s
LoRA rank
64
Fine-tune
$1.40
Serve
$5.00
Lightricks

LTX-2

Lightricks/LTX-2

DiT · native audio + video

Context
Up to 4K · 10s
LoRA rank
32
Fine-tune
$1.60
Serve
$5.50
Wan-AI

Wan 2.2 T2V

Wan-AI/Wan2.2-T2V-A14B

MoE DiT · 27B params · 14B active

Context
720p · 5s
LoRA rank
32
Fine-tune
$1.80
Serve
$6.50

Audio

3
nvidia

Parakeet TDT 1.1B

nvidia/parakeet-tdt-1.1b

ASR · 1.1B params

Context
Streaming
LoRA rank
64
Fine-tune
$0.30
Serve
$0.20
openai

Whisper large-v3

openai/whisper-large-v3

ASR · 1.5B params

Context
30s windows
LoRA rank
64
Fine-tune
$0.60
Serve
$0.40
mistralai

Voxtral 4B TTS

mistralai/Voxtral-4B-TTS-2603

TTS · 4B params

Context
9 languages
LoRA rank
64
Fine-tune
$0.90
Serve
$4.50