Agentic chat completion data preparation
The example on this page is an assistant that answers by calling tools and reading their results: it calls, sees what came back, and either calls again or answers in text.
The shared rules for the input directory, row shapes and validation are in
Overview. This page covers what’s specific to agentic chat
completion. The task extends Chat completion: tool calls
can be followed by tool messages carrying the results of calls, followed by more assistant turns
with or without tool calls.
Job description
Section titled “Job description”Same fields as chat completion, with one difference: tools is required here, and it can’t
be empty. An agentic loop without tools has no results to feed back, so a tool-free task belongs
in plain chat completion.
{
"task_description": "You are a research assistant. Answer the user's question by calling the available tools, reading their results, and replying with a concrete answer grounded in what came back.",
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Look up the current weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "City to look up the weather for"
}
},
"required": ["location"]
}
}
},
{
"type": "function",
"function": {
"name": "search",
"description": "Search for information",
"parameters": {
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "Search query"
}
},
"required": ["query"]
}
}
}
]
}
Training data
Section titled “Training data”Each row is one whole conversation in a single messages array. Three roles:
userturns carry the request, incontent.assistantturns carrycontent,tool_calls, or both, exactly as in chat completion.toolturns carry the result of a call:{"role": "tool", "tool_call_id": ..., "content": ...}. They may only follow an assistant turn that actually made tool calls, and after one, the conversation can continue with anothertoolresult, anassistantturn, or the nextuserturn.
A conversation starts with a user message and ends with an assistant message — typically the
final text answer after the tool results. No system messages; the task_description plays
that part.
The tool_call_id on a result points at the id of the call it answers. A single call and a
single result are matched automatically; with parallel calls the ids must match. Results are
optional per call: a turn can make two calls and carry one result. The examples below set ids on
every call.
Written out, one conversation looks like this:
[
{"role": "user", "content": "What's the weather in Paris?"},
{"role": "assistant", "content": "Let me check that for you.", "tool_calls": [
{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Paris"}}}
]},
{"role": "tool", "tool_call_id": "call_1", "content": "18C, light rain"},
{"role": "assistant", "content": "It's 18C in Paris with light rain."}
]
In train.jsonl that whole array goes on one line. Cover the loop shapes you need: one call
then an answer, chained calls where the second depends on the first’s result, parallel calls,
and follow-up user turns after an answered round. Aim for 20+ conversations.
{"messages": [{"role": "user", "content": "What's the weather in Paris?"}, {"role": "assistant", "content": "Let me check that for you.", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Paris"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "18C, light rain"}, {"role": "assistant", "content": "It's 18C in Paris with light rain."}]}
{"messages": [{"role": "user", "content": "Look up the tallest building, then tell me the weather there."}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "search", "arguments": {"query": "tallest building in the world"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "Burj Khalifa, Dubai"}, {"role": "assistant", "tool_calls": [{"id": "call_2", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Dubai"}}}]}, {"role": "tool", "tool_call_id": "call_2", "content": "34C, sunny"}, {"role": "assistant", "content": "It's the Burj Khalifa in Dubai, where it's 34C and sunny."}]}
{"messages": [{"role": "user", "content": "Compare the weather in Rome and Madrid."}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Rome"}}}, {"id": "call_2", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Madrid"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "22C, clear"}, {"role": "tool", "tool_call_id": "call_2", "content": "27C, sunny"}, {"role": "assistant", "content": "Madrid is warmer at 27C and sunny; Rome is 22C and clear."}]}
A conversation can be split into one example per assistant turn, each predicted from everything before it, tool results included. That matches how the deployed model runs: your code executes the calls, appends the results, and calls the model again (Conversation expansion).
As in chat completion, parallel calls are valid data, but generation defaults to one call per
turn; raise synthgen.max_tool_calls_per_turn if your task needs more.
Test data
Section titled “Test data”Same format as train.jsonl, held out for evaluation. Cover each tool, a multi-step chain, and
a final-answer turn. The main score is llm-as-a-judge-reference-free, which judges the text the
model writes after the tool results, alongside the tool-call metrics against the reference
(Metrics).
{"messages": [{"role": "user", "content": "Is it warm enough for the beach in Barcelona?"}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Barcelona"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "29C, sunny"}, {"role": "assistant", "content": "Yes: 29C and sunny in Barcelona, a good beach day."}]}
{"messages": [{"role": "user", "content": "Where are the next Olympics, and what's the weather there?"}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "search", "arguments": {"query": "next Olympics host city"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "Brisbane, Australia (2032)"}, {"role": "assistant", "tool_calls": [{"id": "call_2", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Brisbane"}}}]}, {"role": "tool", "tool_call_id": "call_2", "content": "24C, partly cloudy"}, {"role": "assistant", "content": "They are in Brisbane, where it's 24C and partly cloudy."}]}
Parallel calls are scored in order, so put the reference calls in the order you want them made.
Config
Section titled “Config”The task type, the two models and the judge:
base:
task: chat-completion-agentic
student_model_name: Qwen3.5-4B
teacher_model_name: zai.glm-5.3-low-thinking
evaluation:
llm_as_a_judge_model_name: zai.glm-5.3-low-thinking
Because tools are always declared here, both models have to be tool-calling capable (Supported models).
synthgen.max_tool_calls_per_turn works as in
chat completion: unset means one call per
generated turn, an integer or unlimited allows more.
To spread the generated data across your tools, add a mutator with one value per tool, as on Chat completion.
Every other field has a default. See Config file for the full table.
Put the files in one directory, here input-dir, and create the seed dataset from it. This
validates your files at the same time:
distil seed-dataset create --data input-dir
The command prints Upload successful. Seed dataset ID: <seed-dataset-id>. Run
teacher evaluation on that ID next.