What Makes a Good Production Trace?
A good trace is one that is about your task and internally consistent — relevant and coherent, in the pipeline’s own terms. It does not have to be correct. The platform scores traces on those two axes, drops the weak ones, and has a teacher committee rewrite the survivors into clean reference answers.
What is a production trace?
A logged interaction between a user and an LLM: the system prompt, the user turns, the assistant turns, and any tool calls and tool results in between. One line of traces.jsonl is one trace, written either as an OpenAI chat completion messages array or as a Langfuse observation object.
Every trace is handled as a multi-turn conversation. A single question-and-answer exchange is just a two-turn conversation, and longer sessions are preserved in full rather than chopped into independent pairs.
What does the platform actually check?
Two scores and a rewrite, with documented defaults you can see and change in the config reference.
| Check | Parameter | Default | What it is asking |
|---|---|---|---|
| Relevance | min_relevance_score |
4 of 5 | Is this conversation about the task in the job description? |
| Coherence | min_coherence_score |
3 of 5 | Does the conversation hang together internally? |
| Relabelling | relabel |
true |
Can a teacher committee rewrite the assistant turns into a cleaner reference? |
| Committee | relabelling_committee_models |
[] |
Which models produce candidate rewrites for the teacher to aggregate |
| System prompt | remove_system_prompt_from_traces |
true |
Strip leading system messages, since the job description carries that content |
The asymmetry between the two thresholds is the interesting part. Relevance is strict at 4 of 5; coherence is deliberately looser at 3 of 5, because the documentation is explicit that lower values “allow more corrupted traces through for committee repair”. Off-topic is fatal. Messy is fixable.
Why is correctness not on that list?
Because the assistant turns get rewritten anyway. With relabel: true, a committee of teacher models rewrites each seed conversation as a whole, turning the assistant turns into cleaner and more consistent reference answers. Whatever your old model said is treated as a draft, not as ground truth.
This is the single most counter-intuitive thing about trace-based training, and it follows directly from what traces are for. Their job is to carry distributional signal — what your users ask, in their words, in their order. The normative signal, what a correct answer looks like, comes from the job description and the schema. Mixing the two up is how teams end up training a model to reproduce their old system’s mistakes with more confidence. Why every path ends in synthetic data works through that split.
What disqualifies a trace?
In practice, four things, and our own traces benchmark measured what each does when you train on them directly rather than filtering.
| Problem | What it looks like | Direct-training cost measured |
|---|---|---|
| Wrong domain | Hotel-booking traces mixed into a restaurant-booking set | 0.694 vs 0.858 (16.4pp) |
| Schema drift | FindRestaurants and search_restaurants both present |
0.585 vs 0.844 (25.9pp) |
| Noisy labels | Chat where a tool call belonged, and vice versa | 0.721 vs 0.844 (12.3pp) |
| Too few distinct traces | Five conversations for a multi-turn task | 0.649 vs 0.852 (20.3pp) |
Wrong domain is what relevance filtering exists to catch, and it is the only one of the four that is genuinely unfixable — a hotel conversation contains no information about restaurant booking. The other three degrade traces as labels while leaving them perfectly good as seeds, which is exactly why the synthetic column of that table barely moves.
Do bad traces still have value?
Yes, and this is the part worth remembering. Traces beyond the seed budget are not discarded — they become unstructured context that keeps generation on-domain, and they pass through unchanged rather than being filtered and rewritten.
So a log full of half-finished conversations, aborted sessions, and requests your old model fumbled still tells a teacher model what your domain sounds like. It just should not be used to define what a correct answer is. The practical implication: do not hand-filter your export. Upload it and let the pipeline sort seeds from context. How much traffic before traces are useful covers where that boundary falls.
How is a trace different from a training example?
A training example is a messages conversation you are asserting is correct. A trace is a conversation you are asserting happened. Everything else follows from that distinction.
| Trace | Training example | |
|---|---|---|
| You vouch for it | No | Yes |
| Gets rewritten | Yes, if selected as a seed | No |
| Gets scored and possibly dropped | Yes | No |
| Minimum you need | Whatever survives filtering | 20 |
If you find yourself curating traces by hand until you would vouch for them, you have stopped doing trace-based training and started doing the minimal-dataset path — which is fine, and usually faster. See fine-tune with 20 examples and three ways to get training data for the comparison.
The full methodology behind the numbers above, including the scenario construction, is in the benchmarking repository.