Evaluate

Know how good your model really is

Upload a test set, build one in the dashboard, or let Nucleus generate it. Every checkpoint gets a domain score against your criteria and your pass threshold. The best one is named.

ft:triage-v3
Checkpoint 3 of 4 · 675 steps
Scored
94.2domain score
threshold 93.0
Speed
142 tok/s
Cost
$0.31 / 1M tokens
Score per checkpoint
Best · ckpt 3
88.6
91
94.2
92.4
ckpt 1ckpt 2ckpt 3ckpt 4
Pass threshold 93.0 · steps 225, 450, 675, 900

Bring your benchmarks

Upload a domain-specific test set and compare accuracy, speed and cost across checkpoints.

Build in the dashboard

Define the questions, the criteria and the pass threshold in the dashboard, with no code to write.

Auto-generated benchmarks

Nucleus generates domain-specific tests and scores each checkpoint against them.

A before-and-after number

Score the base model on the same set in the same run, and the delta from fine-tuning reads straight off the response.

National Neuroscience Institute logo
National University Hospital logo
Mandai Wildlife Group logo
SingHealth logo
Infocomm Media Development Authority logo
Institute for Adult Learning logo
Ministry of Education logo
Nanyang Polytechnic logo
Institute of Singapore Chartered Accountants logo
Ministry of Health logo
National Neuroscience Institute logo
National University Hospital logo
Mandai Wildlife Group logo
SingHealth logo
Infocomm Media Development Authority logo
Institute for Adult Learning logo
Ministry of Education logo
Nanyang Polytechnic logo
Institute of Singapore Chartered Accountants logo
Ministry of Health logo
National Neuroscience Institute logo
National University Hospital logo
Mandai Wildlife Group logo
SingHealth logo
Infocomm Media Development Authority logo
Institute for Adult Learning logo
Ministry of Education logo
Nanyang Polytechnic logo
Institute of Singapore Chartered Accountants logo
Ministry of Health logo
National Neuroscience Institute logo
National University Hospital logo
Mandai Wildlife Group logo
SingHealth logo
Infocomm Media Development Authority logo
Institute for Adult Learning logo
Ministry of Education logo
Nanyang Polytechnic logo
Institute of Singapore Chartered Accountants logo
Ministry of Health logo
Trusted by leading institutions
Two ways to score

Score against a test set, or against the results you record

Evaluations score a checkpoint against a benchmark set you upload, build or generate. Outcome benchmarks score it against the results your systems already record.

An evaluation runs a checkpoint over your test set and scores the outputs, giving you one number for how well it performs. Set the criteria and the pass threshold once, and every checkpoint from a run is scored against them. Pass compare_to and the base model is scored on the same set in the same run, so the delta from fine-tuning is read straight off the response.

Explore Evaluations
Results

Results from LLMs powered by Nucleus

94%

prediction accuracy on equipment failures

14%

yield shortfall traced to one cause

$73.4K

monthly revenue misattributed

6.2 hrs

average time to resolution

Use cases

Where a measurement changed the decision

Manufacturing

Line 3's next failure, booked into the maintenance window

$177Kdowntime prevented this month
18.3 daysadvance notice on a failure
Marketing

$73.4K a month was being credited to the wrong channel

$73.4Kmonthly revenue misattributed
4.12xtrue ROAS against 2.8x reported
Agriculture

Why Block 4B's rice yield came in 14% under target

14%yield shortfall traced to one cause
Day 45when potassium fell 23%

What's included: Evaluate at a glance

Ways into a test set

  • Upload a JSONL test set
  • Build it in the dashboard
  • Auto-generated by Nucleus
  • Hold out a slice of your training data

Scoring

  • Criteria with a pass threshold
  • A domain score per checkpoint
  • Four metrics: exact_match, rouge, pass_rate, grader
  • Baseline comparison with compare_to
  • Accuracy, speed and cost side by side

Runs and results

  • Asynchronous runs, 202 Accepted with an operation ID
  • Poll the operation or the evaluation until it succeeds
  • score, baseline_score and n on the response
  • Any checkpoint by its nucleus:// path

Outcome benchmarks

  • The system of record connected
  • The outcome defined
  • Every checkpoint scored against real results
  • Drift over time
Pricing

Evaluations billed on training meters

Evaluation runs are metered as token consumption on the same meters as everything else, so the bill stays on three quantities: training tokens, inference tokens and storage.

Evaluations

Billed on training meters

Metered as token consumption, on the same meter your fine-tuning jobs run on.

Storage

$0.10per GB / month

Datasets, checkpoints and exported weights are metered by the GB-hour, so you only pay while you keep them.

Pick your path

Pick your path

Evaluations

Score every checkpoint

Explore Evaluations

Outcome benchmarks

Business results as the test set

Explore Outcome benchmarks

Put a number on your next checkpoint

Read the Evaluations reference, or see how evaluation runs are metered.