Skip to content

Test set from traces

This job takes a traces object and builds an updated traces object that carries a test set made from your traces. It relabels some traces into test rows, can generate synthetic test rows from them, and can score your original production model on the relabeled traces. That score is the one your trained model has to beat. The traces the job uses are removed from the updated trace set, so trace processing over it cannot put the same trace in both train and test.

A test set is usually built once and kept across iterations, because scores measured on different test sets are not comparable (Model iterations).

Run the job on a traces object from an upload or from an inference endpoint:

distil traces expand-test-set --output json <traces-id> | jq -r .id

The command creates a new traces object whose source is test_set_expansion and whose parent is the one you named, and prints its ID, the updated traces ID. Poll its status until it reads JOB_SUCCESS:

distil traces status --output json <updated-traces-id> | jq -r .status

The job log is distil traces logs <updated-traces-id>.

--config and --job-description replace the parent’s files for this run. Start from the parent’s own files. Download them into traces/test-set-1:

distil traces download-metadata -d traces/test-set-1 <traces-id>

Edit traces/test-set-1/config.yaml and pass it back:

distil traces expand-test-set --output json \
  --config traces/test-set-1/config.yaml <traces-id> | jq -r .id

Each file is replaced whole, so a field you leave out takes its default rather than the parent’s value.

The traces_to_test_set section controls the job:

traces_to_test_set:
  num_traces_to_relabel: 200
  num_synthetic_examples: 0
  evaluate_original_model: true
  min_relabelled_examples: 0
  • num_traces_to_relabel (default 200) is the number of traces relabeled into test rows, the size of the test set before filtering. Relabeling, relevance filtering and schema checks drop some traces, so the test set comes out smaller.
  • num_synthetic_examples (default 0, a floor: generation runs in batches) grows the test set beyond your traces. The teacher generates these rows with the synthgen settings, using the next 2 × num_synthetic_examples traces as context. Synthetic rows measure agreement with the teacher rather than with production, so read them with more care.
  • evaluate_original_model (default false) scores your production model on the relabeled traces. That score is the reference point for every later result on this test set.
  • min_relabelled_examples (default 0, best kept at 0) sets the smallest test set the job accepts. With fewer surviving rows, the job fails instead of producing a smaller test set.

How a trace is relabeled (the teacher, the committee, relevance filtering) comes from the trace_processing section, the same settings trace processing uses. The full tables are on Config file.

This job removes the traces it uses, and trace processing needs enough of the rest. The traces object has to hold at least:

num_traces_to_relabel + 2 × num_synthetic_examples + num_traces_as_training_base

That is 400 traces at the defaults: 200 for the test set and 200 for the seed dataset. Both jobs refuse a trace set smaller than trace_processing.num_traces_as_training_base.

The test set decides every result that follows, so review it before trace processing. Download the updated traces object into traces/test-set-1/output, read its metrics, and download the original model’s predictions:

distil traces download -d traces/test-set-1/output <updated-traces-id>
distil traces metrics --output json <updated-traces-id>
distil traces download-predictions <updated-traces-id>

The predictions land in <updated-traces-id>-predictions.jsonl in the current directory. The downloaded test.jsonl holds your curated test rows, then the relabeled traces, then the synthetic rows. Check:

  • Size: the row count against num_traces_to_relabel and num_synthetic_examples. A large loss means filtering or schema checks dropped traces, and the job log has the cause.
  • Correctness: the relabeled answers are correct against your job description, and better than the original answers where they differ.
  • Coverage: the test rows cover the topics, lengths and classes your production traffic has.
  • The original model’s score: base_model_performance in traces metrics, on the task’s primary metric, llm-as-a-judge-reference-free (accuracy for classification), on the rows this run relabeled only. The predictions file shows where the original model was marked wrong.

To change the test set, run this job again with different settings. You can also edit the downloaded test.jsonl and upload the directory as a new traces object, which goes to trace processing directly. An edited test set has no original model score: the job scores the original model only on traces it relabels in the same run, and the earlier score was measured on rows that have changed.

One prepared_traces_with_expanded_test_set_post credit per run. A new account starts with 20.

Process the remaining traces into a seed dataset, using the updated traces ID.