Model training
Training fine-tunes your student model on the generated dataset, then evaluates both the untrained and the trained student on your test set, which gives you the comparison with the base student and the teacher.
Before you start
Section titled “Before you start”An empty train.jsonl stops the run, so run
synthetic data generation first. An empty test.jsonl produces a
model with no scores.
The settings that matter are in base and tuning (Config file):
base.student_model_name. The model you’re training. Start withQwen3.5-4B, and go smaller only when the deployment target needs it (Supported models).Qwen3.8-27Btrains only withtuning.use_qlora: true.tuning.per_device_train_batch_size(default 1).tuning.num_train_epochs(default 4).
For a reasoning student, see Reasoning models.
Smoke run: check memory
Section titled “Smoke run: check memory”A smoke run checks that a student fits in memory before the full run spends a training credit. It trains one epoch on the 128 longest training rows and the 32 longest test rows (chosen by length, not per class), with the rest of your config unchanged. Run one per student, at the batch size the full run will use, since students of different sizes fit differently.
Download the training dataset’s config into sweep. It can be read once the generation job
reports JOB_SUCCESS:
distil training-dataset download-metadata -d sweep <training-dataset-id>
Copy sweep/config.yaml once per student, for example to sweep/student-qwen3.5-4b.yaml, and set
base.student_model_name in each copy. Submit a smoke run with each copy. The command prints the
new SLM’s id:
distil slm create-from-training-dataset --smoke --output json \
--config sweep/student-qwen3.5-4b.yaml <training-dataset-id> | jq -r .id
A smoke shows whether the run fits in memory, whether the training loss decreases in the log, and whether base and tuned metrics come back. Its scores come from one epoch on a small subset and say nothing about the full run. When it runs out of memory, Dealing with out-of-memory failures lists what to change.
Each smoke spends a slms_from_training_datasets_smoke_post credit, never a full training
credit.
Full run
Section titled “Full run”The full run uses all the training data and the default 4 epochs, with the settings the smoke fit in memory with. Submit it with the same config file; the command prints the new SLM’s id:
distil slm create-from-training-dataset --output json \
--config sweep/student-qwen3.5-4b.yaml <training-dataset-id> | jq -r .id
Check its status with that id until it reads JOB_SUCCESS:
distil slm status --output json <slm-id> | jq -r .status
Each run spends one slms_from_training_datasets_post credit, and new accounts start with two.
A run takes about 90 minutes. To compare students, submit one run per student config against the
same dataset; the submissions run concurrently and nothing is regenerated.
Each config file replaces the dataset’s config whole, and anything you leave out takes its default, so take a full copy of the downloaded config for each run (How the platform works).
Reading the results
Section titled “Reading the results”Read the scores of the base and the tuned student:
distil slm metrics --output json <slm-id> \
| jq '{base: .base_model_performance, tuned: .tuned_model_performance}'
Download the per-example predictions:
distil slm download-predictions <slm-id>
This writes <slm-id>-slm-predictions.jsonl to the current directory. It holds the tuned model’s
predictions only; the base model’s scores are in the metrics, and its per-example predictions are
not available.
The primary metric is llm-as-a-judge-reference-free for question answering and chat
completion, and accuracy for classification (Metrics). Three numbers
are read together:
baseis the untrained student, the floor.tunedis your trained student.- Teacher is the ceiling, from teacher evaluation.
How much of the base-to-teacher gap the student closed is one number:
closed = (tuned - base) / (teacher - base)
closed near 1 means the student does the task about as well as the teacher; near 0 means
training added little over the base student.
If the test set came from your traces, the original production model’s score is a fourth number:
the model the student would replace. It is in distil traces metrics
(Test set from traces). Scores vary between runs of the same
evaluation (Metrics).
Dealing with out-of-memory failures
Section titled “Dealing with out-of-memory failures”Apply these in order. Each costs more speed or quality than the one before.
- Lower
per_device_train_batch_size. Halve it, down to 1. memory_optimized_training: true. Activation offloading and gradient checkpointing. Much slower.use_qlora: truetogether withmemory_optimized_training: true. A 4-bit base model, about 3x less GPU memory.- Remove the longest ~1% of training rows, which drive peak memory. Download the dataset,
remove the rows from
train.jsonl, and create a new training dataset from the directory withdistil training-dataset create --data. This needs atraining_datasets_download_getcredit, which new accounts don’t have, and atraining_datasets_postcredit.
The first three are config overrides; only the fourth changes the data.
To find the cause of a failed run, print its log:
distil slm logs --output json <slm-id> | jq -r .logs
Search the log for the first error, not the last one. An out-of-memory crash often ends with a
second, unrelated error, thousands of characters after the OutOfMemoryError that caused it.
What can be changed
Section titled “What can be changed”All of these are config overrides on the same training dataset, and nothing is regenerated:
- The student (
base.student_model_name). A larger student has more capacity to learn the same data, at a higher serving cost. A different family can suit a task better at the same size. - How much it trains (
tuning.num_train_epochs,tuning.learning_rate,tuning.per_device_train_batch_size). More optimizer steps fit the training data more closely; at a fixed epoch count, a larger batch means fewer steps. - The LoRA capacity (
tuning.lora_r,tuning.lora_alpha_multiplier). How much of the model training can change. Deployment accepts only somelora_rvalues (Config file). - Evaluation length (
tuning.max_completion_length). How long an answer may be before it is cut off in evaluation.
Serve it, either behind an inference endpoint or on your own machine. To change something and train again, see Model iterations.