Agentic chat completion data preparation
The example on this page is an assistant that answers by calling tools and reading their results: it calls, sees what came back, and either calls again or answers in text.
The shared rules for the input directory, row shapes and validation are in
Overview. This page covers what’s specific to agentic chat
completion. The task extends Chat completion: tool calls can be followed by tool
messages carrying the results of calls, followed by more assistant turns with or without tool calls.
Job description
Section titled “Job description”Same fields as chat completion, with one difference: tools is required here, and it can’t
be empty. An agentic loop without tools has no results to feed back, so a tool-free task belongs
in plain chat completion.
{
"task_description": "You are a research assistant. Answer the user's question by calling the available tools, reading their results, and replying with a concrete answer grounded in what came back.",
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Look up the current weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "City to look up the weather for"
}
},
"required": ["location"]
}
}
},
{
"type": "function",
"function": {
"name": "search",
"description": "Search for information",
"parameters": {
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "Search query"
}
},
"required": ["query"]
}
}
}
]
}
Training data
Section titled “Training data”Each row is one whole conversation in a single messages array. Three roles:
userturns carry the request, incontent.assistantturns carrycontent,tool_calls, or both, exactly as in chat completion.toolturns carry the result of a call:{"role": "tool", "tool_call_id": ..., "content": ...}. They may only follow an assistant turn that actually made tool calls, and after one, the conversation can continue with anothertoolresult, anassistantturn, or the nextuserturn.
A conversation starts with a user message and ends with an assistant message — typically the
final text answer after the tool round-trips. No system messages; the task_description plays
that part.
The tool_call_id on a result points at the id of the call it answers. With a single call and
a single result the pairing is fixed up automatically, but the moment a turn makes parallel
calls, the ids have to match. Results are optional per call: a turn can make two calls and carry
one result. Simplest is to always set ids, as the examples below do.
Written out, one conversation looks like this:
[
{"role": "user", "content": "What's the weather in Paris?"},
{"role": "assistant", "content": "Let me check that for you.", "tool_calls": [
{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Paris"}}}
]},
{"role": "tool", "tool_call_id": "call_1", "content": "18C, light rain"},
{"role": "assistant", "content": "It's 18C in Paris with light rain."}
]
In train.jsonl that whole array goes on one line. Cover the loop shapes you need: one call
then an answer, chained calls where the second depends on the first’s result, parallel calls,
and follow-up user turns after an answered round. Aim for 20+ conversations.
{"messages": [{"role": "user", "content": "What's the weather in Paris?"}, {"role": "assistant", "content": "Let me check that for you.", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Paris"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "18C, light rain"}, {"role": "assistant", "content": "It's 18C in Paris with light rain."}]}
{"messages": [{"role": "user", "content": "Look up the tallest building, then tell me the weather there."}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "search", "arguments": {"query": "tallest building in the world"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "Burj Khalifa, Dubai"}, {"role": "assistant", "tool_calls": [{"id": "call_2", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Dubai"}}}]}, {"role": "tool", "tool_call_id": "call_2", "content": "34C, sunny"}, {"role": "assistant", "content": "The tallest building is the Burj Khalifa in Dubai, where it's currently 34C and sunny."}]}
{"messages": [{"role": "user", "content": "Compare the weather in Rome and Madrid."}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Rome"}}}, {"id": "call_2", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Madrid"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "22C, clear"}, {"role": "tool", "tool_call_id": "call_2", "content": "27C, sunny"}, {"role": "assistant", "content": "Madrid is warmer at 27C and sunny; Rome is 22C and clear."}]}
The model learns every assistant turn, not just the last one: training expands each conversation into one example per assistant turn, predicted from everything before it, tool results included. That’s exactly the loop the deployed model runs, where your code executes the calls, appends the results, and invokes the model again.
As in chat completion, parallel calls are valid data, but generation defaults to one call per
turn; raise synthgen.max_tool_calls_per_turn if your task needs more.
Test data
Section titled “Test data”Same format as train.jsonl, held out for evaluation. Cover each tool, a multi-step chain, and
a final-answer turn, since the text the model writes after the tool results is scored by the LLM
judge just like the calls are scored against the reference.
{"messages": [{"role": "user", "content": "Is it warm enough for the beach in Barcelona?"}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Barcelona"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "29C, sunny"}, {"role": "assistant", "content": "Yes: 29C and sunny in Barcelona, a good beach day."}]}
{"messages": [{"role": "user", "content": "Where are the next Olympics, and what's the weather there?"}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "search", "arguments": {"query": "next Olympics host city"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "Brisbane, Australia (2032)"}, {"role": "assistant", "tool_calls": [{"id": "call_2", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Brisbane"}}}]}, {"role": "tool", "tool_call_id": "call_2", "content": "24C, partly cloudy"}, {"role": "assistant", "content": "The next Summer Olympics are in Brisbane, Australia, where it's currently 24C and partly cloudy."}]}
Evaluation expands each conversation into one line per assistant turn, so the test-set size reported back to you won’t match the number of rows you uploaded. Parallel calls are scored in order, so put the reference calls in the order you want them made.
Config
Section titled “Config”The task type, plus the two models:
base:
task: chat-completion-agentic
student_model_name: Qwen3-1.7B
teacher_model_name: openai.gpt-oss-120b
Because tools are always declared here, the tool-calling model restrictions always apply: students are limited to the Qwen3, Qwen3.5, Llama 3-family, LFM2/LFM2.5, FunctionGemma and Gemma 4 families, and teachers to those marked in the tool-calling column of Supported models.
synthgen.max_tool_calls_per_turn works as in
chat completion: unset means one call per
generated turn, an integer or unlimited allows more.
Every other field has a default. See Config file for the full table.
Create the seed dataset, which validates your files at the same time:
distil seed-dataset create --data ./your-data-dir
Then run teacher evaluation.