Skip to content

Agentic chat completion data preparation

The example on this page is an assistant that answers by calling tools and reading their results: it calls, sees what came back, and either calls again or answers in text.

The shared rules for the input directory, row shapes and validation are in Overview. This page covers what’s specific to agentic chat completion. The task extends Chat completion: tool calls can be followed by tool messages carrying the results of calls, followed by more assistant turns with or without tool calls.

Same fields as chat completion, with one difference: tools is required here, and it can’t be empty. An agentic loop without tools has no results to feed back, so a tool-free task belongs in plain chat completion.

{
  "task_description": "You are a research assistant. Answer the user's question by calling the available tools, reading their results, and replying with a concrete answer grounded in what came back.",
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Look up the current weather for a location",
        "parameters": {
          "type": "object",
          "properties": {
            "location": {
              "type": "string",
              "description": "City to look up the weather for"
            }
          },
          "required": ["location"]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "search",
        "description": "Search for information",
        "parameters": {
          "type": "object",
          "properties": {
            "query": {
              "type": "string",
              "description": "Search query"
            }
          },
          "required": ["query"]
        }
      }
    }
  ]
}

Each row is one whole conversation in a single messages array. Three roles:

  • user turns carry the request, in content.
  • assistant turns carry content, tool_calls, or both, exactly as in chat completion.
  • tool turns carry the result of a call: {"role": "tool", "tool_call_id": ..., "content": ...}. They may only follow an assistant turn that actually made tool calls, and after one, the conversation can continue with another tool result, an assistant turn, or the next user turn.

A conversation starts with a user message and ends with an assistant message — typically the final text answer after the tool round-trips. No system messages; the task_description plays that part.

The tool_call_id on a result points at the id of the call it answers. With a single call and a single result the pairing is fixed up automatically, but the moment a turn makes parallel calls, the ids have to match. Results are optional per call: a turn can make two calls and carry one result. Simplest is to always set ids, as the examples below do.

Written out, one conversation looks like this:

[
  {"role": "user", "content": "What's the weather in Paris?"},
  {"role": "assistant", "content": "Let me check that for you.", "tool_calls": [
    {"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Paris"}}}
  ]},
  {"role": "tool", "tool_call_id": "call_1", "content": "18C, light rain"},
  {"role": "assistant", "content": "It's 18C in Paris with light rain."}
]

In train.jsonl that whole array goes on one line. Cover the loop shapes you need: one call then an answer, chained calls where the second depends on the first’s result, parallel calls, and follow-up user turns after an answered round. Aim for 20+ conversations.

{"messages": [{"role": "user", "content": "What's the weather in Paris?"}, {"role": "assistant", "content": "Let me check that for you.", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Paris"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "18C, light rain"}, {"role": "assistant", "content": "It's 18C in Paris with light rain."}]}
{"messages": [{"role": "user", "content": "Look up the tallest building, then tell me the weather there."}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "search", "arguments": {"query": "tallest building in the world"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "Burj Khalifa, Dubai"}, {"role": "assistant", "tool_calls": [{"id": "call_2", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Dubai"}}}]}, {"role": "tool", "tool_call_id": "call_2", "content": "34C, sunny"}, {"role": "assistant", "content": "The tallest building is the Burj Khalifa in Dubai, where it's currently 34C and sunny."}]}
{"messages": [{"role": "user", "content": "Compare the weather in Rome and Madrid."}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Rome"}}}, {"id": "call_2", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Madrid"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "22C, clear"}, {"role": "tool", "tool_call_id": "call_2", "content": "27C, sunny"}, {"role": "assistant", "content": "Madrid is warmer at 27C and sunny; Rome is 22C and clear."}]}

The model learns every assistant turn, not just the last one: training expands each conversation into one example per assistant turn, predicted from everything before it, tool results included. That’s exactly the loop the deployed model runs, where your code executes the calls, appends the results, and invokes the model again.

As in chat completion, parallel calls are valid data, but generation defaults to one call per turn; raise synthgen.max_tool_calls_per_turn if your task needs more.

Same format as train.jsonl, held out for evaluation. Cover each tool, a multi-step chain, and a final-answer turn, since the text the model writes after the tool results is scored by the LLM judge just like the calls are scored against the reference.

{"messages": [{"role": "user", "content": "Is it warm enough for the beach in Barcelona?"}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Barcelona"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "29C, sunny"}, {"role": "assistant", "content": "Yes: 29C and sunny in Barcelona, a good beach day."}]}
{"messages": [{"role": "user", "content": "Where are the next Olympics, and what's the weather there?"}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "search", "arguments": {"query": "next Olympics host city"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "Brisbane, Australia (2032)"}, {"role": "assistant", "tool_calls": [{"id": "call_2", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Brisbane"}}}]}, {"role": "tool", "tool_call_id": "call_2", "content": "24C, partly cloudy"}, {"role": "assistant", "content": "The next Summer Olympics are in Brisbane, Australia, where it's currently 24C and partly cloudy."}]}

Evaluation expands each conversation into one line per assistant turn, so the test-set size reported back to you won’t match the number of rows you uploaded. Parallel calls are scored in order, so put the reference calls in the order you want them made.

The task type, plus the two models:

base:
  task: chat-completion-agentic
  student_model_name: Qwen3-1.7B
  teacher_model_name: openai.gpt-oss-120b

Because tools are always declared here, the tool-calling model restrictions always apply: students are limited to the Qwen3, Qwen3.5, Llama 3-family, LFM2/LFM2.5, FunctionGemma and Gemma 4 families, and teachers to those marked in the tool-calling column of Supported models.

synthgen.max_tool_calls_per_turn works as in chat completion: unset means one call per generated turn, an integer or unlimited allows more.

Every other field has a default. See Config file for the full table.

Create the seed dataset, which validates your files at the same time:

distil seed-dataset create --data ./your-data-dir

Then run teacher evaluation.