Evaluate · Outcome benchmarks

Benchmark your model on the outcomes your business already records

The results your CRM, ticketing or ERP already record become the test set. Every checkpoint is scored against them, so the score reads as a business outcome.

ft:triage-v3Qwen3.8 27B · LoRA rank 16 · support-eval.jsonl · exact_match
Workflow run
94.2Domain score, checkpoint 3 of 4
1,200 recorded cases
System of record
Tickets
Outcome
Resolution rate
Pass
≥ 93.0
88.6
91
94.2
92.4
ckpt 1ckpt 2ckpt 3ckpt 4
Pass threshold 93.0

The test set is your history

Each row is a real case with the result your system wrote down for it.

The score is a field you already track

Choose the column that holds the outcome and the score is the share of cases the model got right.

Every checkpoint, against your threshold

Each checkpoint of a run is scored beside the base model, and the best one is named.

How it works

From a system of record to a score

Four steps, each one a thing you can see in the dashboard or read off the API.

Point Nucleus at the CRM, the ticketing queue or the ERP where the result is written down. Each case and the result recorded for it become one row of the test set. The rows are held as an evaluation file, the same file type every evaluation reads.

Explore the files reference
Results

Results from LLMs powered by Nucleus

$73.4K

monthly revenue misattributed · Ad performance tracking, marketing

$177K

downtime prevented this month · Predictive maintenance, manufacturing

6.2 hrs

average time to resolution · Critical incident response, IT operations

$2.4K

daily fuel saved across the fleet · Fleet optimisation, fleet operations

Workflows

Where the outcome already lives in a system

IT operations

Every incident past its four-hour SLA, escalated by name

8incidents past their four-hour SLA
6.2 hrsaverage time to resolution
Manufacturing

Line 3's next failure, booked into the maintenance window

$177Kdowntime prevented this month
18.3 daysadvance notice on a failure
Marketing

$73.4K a month was being credited to the wrong channel

$73.4Kmonthly revenue misattributed
4.12xtrue ROAS against 2.8x reported
At a glance

Outcome benchmarks at a glance

Sources

  • JSONL evaluation files, uploaded with purpose=evaluation
  • A held-out slice of the rows you trained on
  • CRM, ticketing and ERP as the system of record

Scoring

  • Metrics exact_match, rouge, pass_rate and grader
  • Baseline comparison with compare_to, baseline_score beside score
  • A pass threshold you set
  • n, the rows scored, on every result

Checkpoints

  • Every checkpoint of a run scored
  • Any nucleus:// checkpoint path as the model
  • The best checkpoint named
  • Runs asynchronously, 202 Accepted, poll the operation

Over time

  • Score again on new records
  • Releases and retrains on one chart
  • Results in the dashboard or through the API
Pricing

Evaluations billed on training meters

Evaluations, billed on training meters

$0.44per 1M training tokens

Qwen3.8 27B. Evaluation runs are metered as token consumption on the same meters as everything else.

Pick your path

More from Evaluate

Evaluations

Score every checkpoint

Explore Evaluations

Score your model on the results you already record

Bring the system that holds the outcome. We will connect it and show you the score.