Chat completion data preparation
The example on this page is a support assistant for a weather service: it chats with the user in free text and calls a tool when one helps answer the request.
The shared rules for the input directory, row shapes and validation are in Overview. This page covers what’s specific to chat completion.
Chat completion is the conversational task without tool results: the model sees user and assistant turns only. If your conversations feed tool outputs back to the model, use Agentic chat completion instead.
Job description
Section titled “Job description”What you expect the model to do, in the words you’d use to prompt an LLM:
task_description. The task itself. There is no system-prompt field anywhere in the input: this text becomes the system prompt the model is trained and evaluated with, so write it with that care.tools. Optional. Every tool the model can call, in OpenAI function-calling format, with unique names. Leave it out for a pure chat model with no tools. If your data contains a tool call, the tool has to be declared here, and every call validates against these schemas.llm_as_a_judge_instructions. Optional. Predictions are free text scored by an LLM judge, so this field does real work here.
{
"task_description": "You are a friendly assistant for a weather service. Chat with the user, and call a tool when it helps answer the request. Answer briefly and concretely, and stay on the topics of weather and travel planning.",
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Look up the current weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "City to look up the weather for"
}
},
"required": ["location"]
}
}
},
{
"type": "function",
"function": {
"name": "search",
"description": "Search for information",
"parameters": {
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "Search query"
}
},
"required": ["query"]
}
}
}
]
}
Training data
Section titled “Training data”Each row is one whole conversation in a single messages array. Two roles only:
userturns carry the request, incontent.assistantturns carrycontent,tool_calls, or both. A turn with neither fails validation.
A conversation starts with a user message, ends with an assistant message, and the two roles
alternate strictly. No system messages (the task description plays that part), and no tool
messages; tool results are what separates this task from
Agentic chat completion.
Written out, one conversation looks like this:
[
{"role": "user", "content": "What's the weather in Paris?"},
{"role": "assistant", "content": "Let me check that for you.", "tool_calls": [
{"type": "function", "function": {"name": "get_weather", "arguments": {"location": "Paris"}}}
]},
{"role": "user", "content": "And is there anything worth seeing there this weekend?"},
{"role": "assistant", "tool_calls": [
{"type": "function", "function": {"name": "search", "arguments": {"query": "Paris events this weekend"}}}
]}
]
In tool_calls, arguments is a JSON object, not a JSON-encoded string. The tool schemas in
job_description.json use the OpenAI function-calling format.
In train.jsonl that whole array goes on one line. Mix the shapes your production conversations
take: plain chat turns, turns that only call a tool, and turns that talk and call at once. Aim
for 20+ conversations.
{"messages": [{"role": "user", "content": "What's the weather in Paris?"}, {"role": "assistant", "content": "Let me check that for you.", "tool_calls": [{"type": "function", "function": {"name": "get_weather", "arguments": {"location": "Paris"}}}]}]}
{"messages": [{"role": "user", "content": "Hi! Can you help me plan a trip?"}, {"role": "assistant", "content": "Of course. Where are you headed, and when?"}, {"role": "user", "content": "Lisbon, next week."}, {"role": "assistant", "content": "Let me check the weather there.", "tool_calls": [{"type": "function", "function": {"name": "get_weather", "arguments": {"location": "Lisbon"}}}]}]}
{"messages": [{"role": "user", "content": "Thanks, that's all I needed."}, {"role": "assistant", "content": "Happy to help. Ask me any time."}]}
{"messages": [{"role": "user", "content": "Compare the weather in Rome and Madrid for me."}, {"role": "assistant", "content": "I'll check both.", "tool_calls": [{"type": "function", "function": {"name": "get_weather", "arguments": {"location": "Rome"}}}, {"type": "function", "function": {"name": "get_weather", "arguments": {"location": "Madrid"}}}]}]}
The last row makes two calls in one turn. Parallel calls are valid data, but by default the
teacher generates at most one call per assistant turn; raise
synthgen.max_tool_calls_per_turn if your task needs more (see Config below).
Test data
Section titled “Test data”Same format as train.jsonl, held out for evaluation. This is what every score you’ll see is
measured against: the teacher evaluation, the base student and the tuned student.
Cover the shapes that matter to you: turns that should chat, turns that should call a tool, and follow-ups that depend on earlier turns.
{"messages": [{"role": "user", "content": "Do I need an umbrella in London today?"}, {"role": "assistant", "content": "Let me look at the current weather.", "tool_calls": [{"type": "function", "function": {"name": "get_weather", "arguments": {"location": "London"}}}]}]}
{"messages": [{"role": "user", "content": "What can you do?"}, {"role": "assistant", "content": "I can check the current weather for any city and search for travel information. What do you need?"}]}
{"messages": [{"role": "user", "content": "Find me something to do in Berlin."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "search", "arguments": {"query": "things to do in Berlin"}}}]}]}
A conversation can be split into one example per assistant turn for training and evaluation, so the number of scored examples can exceed the number of rows you uploaded (Conversation expansion).
The main score is llm-as-a-judge-reference-free, alongside the tool-call metrics
(Metrics). Parallel calls are scored in order: a turn with the right
calls in the wrong order counts as a miss, so put the reference calls in the order you want them
made.
Config
Section titled “Config”The task type, the two models and the judge:
base:
task: chat-completion
student_model_name: Qwen3.5-4B
teacher_model_name: zai.glm-5.3-low-thinking
evaluation:
llm_as_a_judge_model_name: zai.glm-5.3-low-thinking
If the job description declares tools, both models have to be tool-calling capable (Supported models). With no tools declared, any supported model works.
One parameter is specific to this task family:
synthgen:
max_tool_calls_per_turn: unlimited
It caps how many tool calls a generated assistant turn may contain. Unset, the limit is 1 when
tools are declared and 0 when they aren’t. Set an integer or unlimited to allow parallel
calls, or 0 to forbid calls outright. Declaring no tools while setting it above zero fails
validation, as does declaring tools while setting it to 0.
Every other field has a default. See Config file for the full table and Supported models for the values you can use.
Synthetic data with tools
Section titled “Synthetic data with tools”Synthetic data generation does not target tools one at a time: every generation call sees all the declared tools, and the teacher decides which tool each example uses, so some tools can end up rare or missing. When the job description declares tools, add a mutator with one value per tool, so every generation call is asked for a specific one:
synthgen:
mutators:
- name: tool
values:
- "the conversation calls get_weather"
- "the conversation calls search"
- "the assistant answers in text with no tool call"
The last value covers turns that should answer without a tool; leave it out if every turn should
call one. Set target_distribution: match_seed to follow the tool mix of your training rows
instead of an even split, or give one weight per value. Count the tools in the output of a
smoke run before the full run
(Synthetic data generation).
Put the files in one directory, here input-dir, and create the seed dataset from it. This
validates your files at the same time:
distil seed-dataset create --data input-dir
The command prints Upload successful. Seed dataset ID: <seed-dataset-id>. Run
teacher evaluation on that ID next.