EVALUATE · EVALUATIONS

Score every checkpoint, name the best one

Upload a test set, build one in the dashboard, or let Nucleus generate it. Set the criteria and a pass threshold, and one run scores every checkpoint of your fine-tuned model against them.

ft:triage-v3
Checkpoint 3 of 4 · 675 steps
Scored
94.2domain score
threshold 93.0
Speed
142 tok/s
Cost
$0.31 / 1M tokens
Score per checkpoint
Best · ckpt 3
88.6
91
94.2
92.4
ckpt 1ckpt 2ckpt 3ckpt 4
Pass threshold 93.0 · steps 225, 450, 675, 900

Bring your benchmarks

Upload domain-specific test sets and compare accuracy, speed and cost.

Build in the dashboard

Define questions, criteria and pass thresholds, with no code required.

Auto-generated benchmarks

Nucleus generates domain-specific tests and scores each checkpoint.

A before-and-after number

Pass compare_to and the response carries baseline_score beside score, so the gain from fine-tuning reads straight off it.

How a checkpoint is scored

One set, one run, one verdict per checkpoint

A run takes the benchmark set and scores each checkpoint of the fine-tuned model against its criteria. A checkpoint under the threshold is held, and the highest score is named the best. In this run the third checkpoint clears 93.0 and the fourth, the newest, falls back under it.

support-eval.jsonl4 criteria · pass ≥ 93.0
Ready
Metric
grader
exact_match · rouge · pass_rate · grader
Correct routing
400 cases
Resolution rate
300 cases
Tone compliance
300 cases
SLA met
200 cases
Add a criterion
1,200 test cases
Every checkpoint, one runft:triage-v3 · pass ≥ 93.0
ckpt 1225 steps
88.6Held · 4.4 below
ckpt 2450 steps
91.0Held · 2.0 below
ckpt 3675 steps
94.2Best
ckpt 4900 steps
92.4Held · 0.6 below
Explore the evaluations API
Results

Results from LLMs powered by Nucleus

94%

prediction accuracy on equipment failures, predictive maintenance

$73.4K

monthly revenue re-attributed in marketing, ad performance

47 vehicles

rerouted in a single fleet optimisation, fleet optimisation

6.2 hrs

average time to resolution, incident response

USE CASES

Three workflows from the use cases page

Each is an example workflow, powered by Nucleus, and every figure comes from the one it describes.

Procurement

Six vendor contracts ranked before the renewal meeting

6 vendorsunder review for Q4 renewal
3 axescost, quality, response
Manufacturing

Line 3's next failure, booked into the maintenance window

$177Kdowntime prevented this month
18.3 daysadvance notice on a failure
Agriculture

Why Block 4B's rice yield came in 14% under target

14%yield shortfall traced to one cause
Day 45when potassium fell 23%
AT A GLANCE

Everything an evaluation carries

Ways in

  • Upload a JSONL test set, purpose evaluation
  • Build questions, criteria and thresholds in the dashboard
  • Domain benchmarks generated by Nucleus

Scoring

  • Metrics exact_match, rouge, pass_rate and grader
  • A domain score per checkpoint
  • compare_to scores the base model in the same run
  • score, baseline_score and n on the response

Criteria and thresholds

  • A case count per criterion
  • One pass threshold per set
  • A checkpoint under the threshold is held
  • The highest score is named the best

API

  • POST /v1/evaluations answers 202 with an operation
  • GET /v1/evaluations/{id}
  • Poll the operation or the evaluation until succeeded
  • Any checkpoint by its nucleus:// path
PRICING

Evaluations billed on training meters

Evaluation runs are metered as token consumption on the same meters as everything else. There is no separate charge type.

Training tokens

$0.44per 1M training tokens

Fine-tuning jobs and training-loop steps are metered per million tokens processed, whichever way you train.

Pick your path

Also in Evaluate

Outcome benchmarks

Business results as the test set

Explore Outcome benchmarks

Score your next checkpoint

Read the evaluations reference, or see what a run costs on the training meters.