Connect your data, train on what matters
Bring your JSONL or connect a data source. Pick a base model, and managed GPUs handle the rest, from the dashboard or the API. Your first fine-tune runs in minutes.
Connect your sources
SaaS platforms, databases, APIs and documents. Nucleus generates training-ready datasets.
Automated training
Scheduling, GPU allocation and checkpointing, handled. Set up once, run on your schedule.
Full loop access
Forward-backward passes and optimiser steps as API calls. RL, DPO and distillation are yours to write.
One platform from your data to a trained checkpoint
Upload a dataset, submit a job and poll until it finishes. The platform runs the training loop on managed GPUs, so there is no gradient code to write. An evaluation against the base model turns the result into a before-and-after score.
Results from LLMs powered by Nucleus
Monthly revenue misattributed, found
Downtime prevented this month
Daily fuel saved across the fleet
Yield shortfall traced to one cause
Where a trained model went to work
$73.4K a month was being credited to the wrong channel
Line 3's next failure, booked into the maintenance window
47 trucks rerouted while they were still moving
What Train includes
Data
- JSONL chat format, system, user and assistant
- A validation file reports val_loss
- Row-level errors on upload
- Sources: SaaS, databases, APIs, CSV files, documents
- Generated datasets, synced hourly
Managed training
- Supervised (sft) and DPO
- epochs, learning_rate, batch_size, lora_rank
- Events stream with warn events
- Cancel stops the meter
- Idempotency-Key on submit
- Webhook on completion
Training loop
- Training runs on allocated GPUs
- Forward-backward and optim-step
- Cross-entropy, importance sampling, PPO, CISPO, DRO
- Custom-loss exchange
- Metrics history per step
- Checkpoint, resume and load state
- Samplers from a checkpoint
Models
- 25 open bases, six modalities
- A trainable flag per model
- LoRA rank ceiling 16, 32 or 64
- Context windows to 1M
- Models served from base weights when not trainable
Training tokens
Training tokens
Fine-tuning jobs and training-loop steps are metered per million tokens processed, whichever way you train.
Where next in Train
Fine-tuning
Managed training on 25 open models
Training loop
DocsRL, DPO and distillation as API calls
Data connectors
Your systems as training data
Models
Open bases across six modalities
Read on
Managed fine-tuning
Go from a raw JSONL dataset to a served private model with the managed job API.
The training loop
Drive gradients yourself with rendered token IDs, forward-backward passes and optimiser steps.
Mandai Wildlife Group transforms animal care operations with Nucleus
How Mandai Wildlife Group uses Nucleus AI to improve animal care and visitor experiences.
Start with the data you already have
Bring your JSONL or connect a source. Your first fine-tune runs in minutes.