Benchmark your model on the outcomes your business already records
The results your CRM, ticketing or ERP already record become the test set. Every checkpoint is scored against them, so the score reads as a business outcome.
The test set is your history
Each row is a real case with the result your system wrote down for it.
The score is a field you already track
Choose the column that holds the outcome and the score is the share of cases the model got right.
Every checkpoint, against your threshold
Each checkpoint of a run is scored beside the base model, and the best one is named.
From a system of record to a score
Four steps, each one a thing you can see in the dashboard or read off the API.
Point Nucleus at the CRM, the ticketing queue or the ERP where the result is written down. Each case and the result recorded for it become one row of the test set. The rows are held as an evaluation file, the same file type every evaluation reads.
Results from LLMs powered by Nucleus
monthly revenue misattributed · Ad performance tracking, marketing
downtime prevented this month · Predictive maintenance, manufacturing
average time to resolution · Critical incident response, IT operations
daily fuel saved across the fleet · Fleet optimisation, fleet operations
Where the outcome already lives in a system
Every incident past its four-hour SLA, escalated by name
Line 3's next failure, booked into the maintenance window
$73.4K a month was being credited to the wrong channel
Outcome benchmarks at a glance
Sources
- JSONL evaluation files, uploaded with purpose=evaluation
- A held-out slice of the rows you trained on
- CRM, ticketing and ERP as the system of record
Scoring
- Metrics exact_match, rouge, pass_rate and grader
- Baseline comparison with compare_to, baseline_score beside score
- A pass threshold you set
- n, the rows scored, on every result
Checkpoints
- Every checkpoint of a run scored
- Any nucleus:// checkpoint path as the model
- The best checkpoint named
- Runs asynchronously, 202 Accepted, poll the operation
Over time
- Score again on new records
- Releases and retrains on one chart
- Results in the dashboard or through the API
Evaluations billed on training meters
Evaluations, billed on training meters
Qwen3.8 27B. Evaluation runs are metered as token consumption on the same meters as everything else.
More from Evaluate
Evaluations
Score every checkpoint
Read on
Score your model on the results you already record
Bring the system that holds the outcome. We will connect it and show you the score.