DEPLOY · ON-PREMISE DEPLOYMENT

Your fine-tuned LLM, running on your site

Nucleus delivers GPU hardware to your premises, pre-loaded with a model fine-tuned on your data. Your team gets a private inference endpoint on your own network, and nothing ever leaves the building.

On your sitebehind your firewall
running
GPU servers
delivered and installed
Model
fine-tuned on your data
Network
no egress required
Maintenance
Nucleus, under SLA
What landsdelivered, installed, maintained
Rack
2 × H100 servers
Power
2 × 3 kW, redundant
Network
10 GbE to your LAN
Model
pre-loaded, fine-tuned
Support
SLA

Dedicated GPU hardware

Servers sized to your workload, delivered and installed on your premises.

Your model, pre-loaded

Hardware arrives running an LLM fine-tuned on your data, ready on day one.

Fully air-gapped

Inference runs on your LAN with no connection back to Nucleus.

Managed by Nucleus

We install the stack, monitor the hardware, and deliver model updates on site.

How it works

How it runs on your site

Train on your data with managed fine-tuning jobs on Nucleus GPUs. When the checkpoint is right, Nucleus builds the machines with it loaded and installs them in your server room. Your applications then call an endpoint on your LAN.

01Size itthe workload sets the rack
02Trainon your data, on our GPUs
03Deliverinstalled on your site
04Operatemaintained under SLA
Read the managed fine-tuning guide
Results

Results from LLMs powered by Nucleus

$177K

Downtime prevented this month, from Predictive Maintenance

18.3 days

Advance notice on a failure, from Predictive Maintenance

8

Incidents past their four-hour SLA, from Critical Incident Response

Same day

From site reading to client report, from Real-Time Field Data Collection

Use cases

Workflows powered by a Nucleus-trained model

Manufacturing

Line 3's next failure, booked into the maintenance window

$177Kdowntime prevented this month
18.3 daysadvance notice on a failure
IT operations

Every incident past its four-hour SLA, escalated by name

8incidents past their four-hour SLA
6.2 hrsaverage time to resolution
Field engineering

Cable survey readings processed while the crew was still on site

Same dayfrom site reading to client report
9systems on one thread

What comes with the hardware

The hardware

  • Sized to your throughput
  • A single tower to a full rack
  • Delivered, installed and burnt in by Nucleus
  • Scheduled maintenance on site

The model

  • Fine-tuned on your data on Nucleus
  • Pre-loaded before delivery
  • Model updates delivered on site
  • Checkpoints export as a LoRA adapter, merged safetensors or GGUF

The endpoint

  • OpenAI-compatible
  • Reachable only from your network
  • Your existing OpenAI client with a new base URL
  • Health and GPU utilisation from the same endpoint

Security

  • Prompts and completions stay on your network
  • Weights only on hardware you control
  • Runs fully air-gapped when required
  • No runtime connection back to Nucleus
Pricing

Run Nucleus in your own environment

Platform licence

from $50,000/ year

Private and on-prem deployment. Built for data-residency and air-gapped requirements. Hardware is billed separately.

Pick your path

More in Deploy

Harness

Permissions, scoped keys and audit

Explore Harness

Tool use

Your private APIs as skills

Explore Tool use

MCP

Your model in your tools

Explore MCP

RAG

Answers from your documents

Explore RAG

Webhooks

Signed events to your systems

Explore Webhooks

Bring the model to the data

Tell us the workload and the room it has to run in.