Test set from traces
This job takes a traces object and builds an updated traces object that carries a test set made from your traces. It relabels some traces into test rows, can generate synthetic test rows from them, and can score your original production model on the relabeled traces. That score is the one your trained model has to beat. The traces the job uses are removed from the updated trace set, so trace processing over it cannot put the same trace in both train and test.
A test set is usually built once and kept across iterations, because scores measured on different test sets are not comparable (Model iterations).
Running it
Section titled “Running it”Run the job on a traces object from an upload or from an inference endpoint:
distil traces expand-test-set --output json <traces-id> | jq -r .id
The command creates a new traces object whose source is test_set_expansion and whose parent is
the one you named, and prints its ID, the updated traces ID. Poll its status until it reads
JOB_SUCCESS:
distil traces status --output json <updated-traces-id> | jq -r .status
The job log is distil traces logs <updated-traces-id>.
--config and --job-description replace the parent’s files for this run. Start from the
parent’s own files. Download them into traces/test-set-1:
distil traces download-metadata -d traces/test-set-1 <traces-id>
Edit traces/test-set-1/config.yaml and pass it back:
distil traces expand-test-set --output json \
--config traces/test-set-1/config.yaml <traces-id> | jq -r .id
Each file is replaced whole, so a field you leave out takes its default rather than the parent’s value.
The config
Section titled “The config”The traces_to_test_set section controls the job:
traces_to_test_set:
num_traces_to_relabel: 200
num_synthetic_examples: 0
evaluate_original_model: true
min_relabelled_examples: 0
num_traces_to_relabel(default 200) is the number of traces relabeled into test rows, the size of the test set before filtering. Relabeling, relevance filtering and schema checks drop some traces, so the test set comes out smaller.num_synthetic_examples(default 0, a floor: generation runs in batches) grows the test set beyond your traces. The teacher generates these rows with thesynthgensettings, using the next2 × num_synthetic_examplestraces as context. Synthetic rows measure agreement with the teacher rather than with production, so read them with more care.evaluate_original_model(defaultfalse) scores your production model on the relabeled traces. That score is the reference point for every later result on this test set.min_relabelled_examples(default 0, best kept at 0) sets the smallest test set the job accepts. With fewer surviving rows, the job fails instead of producing a smaller test set.
How a trace is relabeled (the teacher, the committee, relevance filtering) comes from the
trace_processing section, the same settings trace processing uses. The full tables are on
Config file.
How many traces you need
Section titled “How many traces you need”This job removes the traces it uses, and trace processing needs enough of the rest. The traces object has to hold at least:
num_traces_to_relabel + 2 × num_synthetic_examples + num_traces_as_training_base
That is 400 traces at the defaults: 200 for the test set and 200 for the seed dataset. Both jobs
refuse a trace set smaller than trace_processing.num_traces_as_training_base.
Review the test set
Section titled “Review the test set”The test set decides every result that follows, so review it before trace processing. Download
the updated traces object into traces/test-set-1/output, read its metrics, and download the
original model’s predictions:
distil traces download -d traces/test-set-1/output <updated-traces-id>
distil traces metrics --output json <updated-traces-id>
distil traces download-predictions <updated-traces-id>
The predictions land in <updated-traces-id>-predictions.jsonl in the current directory. The
downloaded test.jsonl holds your curated test rows, then the relabeled traces, then the synthetic rows.
Check:
- Size: the row count against
num_traces_to_relabelandnum_synthetic_examples. A large loss means filtering or schema checks dropped traces, and the job log has the cause. - Correctness: the relabeled answers are correct against your job description, and better than the original answers where they differ.
- Coverage: the test rows cover the topics, lengths and classes your production traffic has.
- The original model’s score:
base_model_performanceintraces metrics, on the task’s primary metric,llm-as-a-judge-reference-free(accuracyfor classification), on the rows this run relabeled only. The predictions file shows where the original model was marked wrong.
To change the test set, run this job again with different settings. You can also edit the
downloaded test.jsonl and upload the directory as a new traces object, which goes to trace
processing directly. An edited test set has no original model score: the job scores the original
model only on traces it relabels in the same run, and the earlier score was measured on rows
that have changed.
Credits
Section titled “Credits”One prepared_traces_with_expanded_test_set_post credit per run. A new account starts with 20.
Process the remaining traces into a seed dataset, using the updated traces ID.