Traces vs Synthetic Data: Which Trains a Better Model?
Synthetic data generated from traces wins, by up to 25.9 percentage points. On clean curated data the two approaches tie. The moment traces carry realistic noise, training on them directly drops 14 to 28 points below the curated ceiling while the synthetic pipeline stays within 2 points of it.
Which approach wins on the numbers?
Synthetic-from-traces, in every scenario, and by a margin that scales with how messy the traces are.
| Scenario | Synthetic from traces | Direct on traces | Delta |
|---|---|---|---|
| Curated data (human-annotated) | 0.866 | 0.864 | +0.002 (tie) |
| Noisy labels (50% corrupted tool calls) | 0.844 | 0.721 | +12.3pp |
| Schema drift (randomised function names) | 0.844 | 0.585 | +25.9pp |
| Low data (5 traces only) | 0.852 | 0.649 | +20.3pp |
| Irrelevant trace mixing (80% wrong domain) | 0.858 | 0.694 | +16.4pp |
LLM-as-a-judge score on roughly 360 evaluation turn pairs from 34 held-out multi-turn conversations. Best teacher, GLM-5 at 744B parameters: 0.835.
The first row is the control, and it matters as much as the rest. When both pipelines train on clean human-annotated data they land within 0.002 of each other, so the synthetic step isn’t adding quality out of nowhere, and it isn’t degrading anything either. It’s purely a noise-removal layer, which is why it makes no measurable difference on perfect data and a large one on realistic data.
Full write-up: why training on production traces fails.
How was this measured?
On the Schema-Guided Dialogue dataset from Google Research, a corpus of over 20,000 human-curated multi-turn dialogues across 20 domains. The target task was the Restaurants_1 booking agent, with FindRestaurants, ReserveRestaurant, and respond_to_user tools.
After reserving 34 conversations as a shared test set, 327 clean traces remained as the canonical training source. Each scenario applies one deliberate corruption to those 327 and runs both pipelines on the result.
| Parameter | Value |
|---|---|
| Student model | Qwen3-1.7B, LoRA rank 64 |
| Generation teacher | GLM-5 |
| Judge | GPT-OSS-120B |
| Training | 4 epochs, learning rate 5e-5 |
| Synthetic set size | ~2,000 multi-turn conversations per scenario |
Both pipelines shared the test set, student, hyperparameters and evaluation harness. The only variable was whether traces became labels or became context. Code and data: the benchmarking repository.
What are production traces actually good at?
Carrying the shape of your problem. Traces show real user intents, the vocabulary people actually use, realistic conversation flows, and genuine tool-usage patterns. A hand-written seed set reliably captures none of that, because nobody invents the awkward inputs that real users produce daily.
That’s distributional signal, and it’s the reason to bother with traces at all. The benchmark’s low-data row is the clearest demonstration: five real traces, used as seed context, got a 1.7B student to 0.852. Five real conversations carry a lot of information, but almost none of it is usable as labels.
What’s synthetic generation better at?
Everything to do with correctness. The teacher reads the task description and the tool schema alongside the traces, so it knows what should happen, not only what did.
Then a validation layer checks each generated example against the target schema, filters near-duplicates, and rejects length outliers. That layer has no equivalent in direct training, where a malformed trace is simply a malformed training example.
The combination is the whole trick: traces supply the domain, the schema supplies the ground truth, and validation enforces it. Why every path ends in synthetic data covers the mechanism, and fine-tune with synthetic data the practice.
Why did direct training collapse under schema drift?
Because the model had no way to know which name was right. With 21 randomised function names spread across the traces (search_restaurants, lookup_restaurants, find_places_to_eat), direct training simply can’t learn a single correct tool vocabulary, and it scored 0.585.
The synthetic pipeline reads the correct schema from the job description, ignores the wrong names in the traces, and generates examples using the real API. Same input, 0.844.
Schema drift is also the most common of these failure modes in real logs, because APIs get renamed and old versions run alongside new ones for months. If you’ve ever shipped a breaking change to a tool schema, your trace archive has this problem.
Did the student really beat every teacher?
Yes. Five large models were evaluated as teachers on the same test set to establish ceilings.
| Teacher model | LLM-as-a-judge | Std |
|---|---|---|
| GLM-5 | 0.835 | 0.006 |
| Qwen3-235B | 0.768 | 0.018 |
| GPT-OSS-120B | 0.765 | 0.020 |
| MiniMax-M2 | 0.762 | 0.010 |
| DeepSeek-3.2 | 0.744 | 0.014 |
The fine-tuned 1.7B student scored 0.844 to 0.866 across every scenario, above GLM-5’s 0.835, which is a 744B model and therefore roughly 437x larger. Note that this only holds within the task. We’re not claiming a 1.7B model is generally better than GLM-5. See can a small model beat its teacher for why specialisation produces this result.
Which should you pick?
Use traces as seeds and context, and don’t train on them as labels. The only case for direct training is one you almost certainly don’t have.
| Your situation | Use |
|---|---|
| Traces you’d personally vouch for, every one | Either, they tie |
| Traces from a live system, unaudited | Synthetic from traces |
| Multiple prompt or schema versions in the log | Synthetic from traces, emphatically |
| A handful of traces and nothing else | Synthetic from traces |
| No production system at all | Hand-written seeds, see fine-tune with 20 examples |
The first row is the honest exception, and it’s rare. That guarantee is effectively impossible without human curation, at which point you’ve paid the labelling cost anyway.
For multi-turn agents there’s one further consequence. A per-turn error rate compounds: at 95% per-turn accuracy only about 35% of 20-turn interactions come out fully correct. Tool-timing errors are the worst kind, because the output looks syntactically fine. Turn production traces into training data is the procedure.