Skip to content

Model training

Training fine-tunes your student model on the generated dataset, then evaluates both the untrained and the trained student on your test set, which gives you the comparison with the base student and the teacher.

An empty train.jsonl stops the run, so run synthetic data generation first. An empty test.jsonl produces a model with no scores.

The settings that matter are in base and tuning (Config file):

  • base.student_model_name. The model you’re training. Start with Qwen3.5-4B, and go smaller only when the deployment target needs it (Supported models). Qwen3.8-27B trains only with tuning.use_qlora: true.
  • tuning.per_device_train_batch_size (default 1).
  • tuning.num_train_epochs (default 4).

For a reasoning student, see Reasoning models.

A smoke run checks that a student fits in memory before the full run spends a training credit. It trains one epoch on the 128 longest training rows and the 32 longest test rows (chosen by length, not per class), with the rest of your config unchanged. Run one per student, at the batch size the full run will use, since students of different sizes fit differently.

Download the training dataset’s config into sweep. It can be read once the generation job reports JOB_SUCCESS:

distil training-dataset download-metadata -d sweep <training-dataset-id>

Copy sweep/config.yaml once per student, for example to sweep/student-qwen3.5-4b.yaml, and set base.student_model_name in each copy. Submit a smoke run with each copy. The command prints the new SLM’s id:

distil slm create-from-training-dataset --smoke --output json \
  --config sweep/student-qwen3.5-4b.yaml <training-dataset-id> | jq -r .id

A smoke shows whether the run fits in memory, whether the training loss decreases in the log, and whether base and tuned metrics come back. Its scores come from one epoch on a small subset and say nothing about the full run. When it runs out of memory, Dealing with out-of-memory failures lists what to change.

Each smoke spends a slms_from_training_datasets_smoke_post credit, never a full training credit.

The full run uses all the training data and the default 4 epochs, with the settings the smoke fit in memory with. Submit it with the same config file; the command prints the new SLM’s id:

distil slm create-from-training-dataset --output json \
  --config sweep/student-qwen3.5-4b.yaml <training-dataset-id> | jq -r .id

Check its status with that id until it reads JOB_SUCCESS:

distil slm status --output json <slm-id> | jq -r .status

Each run spends one slms_from_training_datasets_post credit, and new accounts start with two. A run takes about 90 minutes. To compare students, submit one run per student config against the same dataset; the submissions run concurrently and nothing is regenerated.

Each config file replaces the dataset’s config whole, and anything you leave out takes its default, so take a full copy of the downloaded config for each run (How the platform works).

Read the scores of the base and the tuned student:

distil slm metrics --output json <slm-id> \
  | jq '{base: .base_model_performance, tuned: .tuned_model_performance}'

Download the per-example predictions:

distil slm download-predictions <slm-id>

This writes <slm-id>-slm-predictions.jsonl to the current directory. It holds the tuned model’s predictions only; the base model’s scores are in the metrics, and its per-example predictions are not available.

The primary metric is llm-as-a-judge-reference-free for question answering and chat completion, and accuracy for classification (Metrics). Three numbers are read together:

  • base is the untrained student, the floor.
  • tuned is your trained student.
  • Teacher is the ceiling, from teacher evaluation.

How much of the base-to-teacher gap the student closed is one number:

closed = (tuned - base) / (teacher - base)

closed near 1 means the student does the task about as well as the teacher; near 0 means training added little over the base student.

If the test set came from your traces, the original production model’s score is a fourth number: the model the student would replace. It is in distil traces metrics (Test set from traces). Scores vary between runs of the same evaluation (Metrics).

Apply these in order. Each costs more speed or quality than the one before.

  1. Lower per_device_train_batch_size. Halve it, down to 1.
  2. memory_optimized_training: true. Activation offloading and gradient checkpointing. Much slower.
  3. use_qlora: true together with memory_optimized_training: true. A 4-bit base model, about 3x less GPU memory.
  4. Remove the longest ~1% of training rows, which drive peak memory. Download the dataset, remove the rows from train.jsonl, and create a new training dataset from the directory with distil training-dataset create --data. This needs a training_datasets_download_get credit, which new accounts don’t have, and a training_datasets_post credit.

The first three are config overrides; only the fourth changes the data.

To find the cause of a failed run, print its log:

distil slm logs --output json <slm-id> | jq -r .logs

Search the log for the first error, not the last one. An out-of-memory crash often ends with a second, unrelated error, thousands of characters after the OutOfMemoryError that caused it.

All of these are config overrides on the same training dataset, and nothing is regenerated:

  • The student (base.student_model_name). A larger student has more capacity to learn the same data, at a higher serving cost. A different family can suit a task better at the same size.
  • How much it trains (tuning.num_train_epochs, tuning.learning_rate, tuning.per_device_train_batch_size). More optimizer steps fit the training data more closely; at a fixed epoch count, a larger batch means fewer steps.
  • The LoRA capacity (tuning.lora_r, tuning.lora_alpha_multiplier). How much of the model training can change. Deployment accepts only some lora_r values (Config file).
  • Evaluation length (tuning.max_completion_length). How long an answer may be before it is cut off in evaluation.

Serve it, either behind an inference endpoint or on your own machine. To change something and train again, see Model iterations.