Skip to content

Agentic chat completion data preparation

The example on this page is an assistant that answers by calling tools and reading their results: it calls, sees what came back, and either calls again or answers in text.

The shared rules for the input directory, row shapes and validation are in Overview. This page covers what’s specific to agentic chat completion. The task extends Chat completion: tool calls can be followed by tool messages carrying the results of calls, followed by more assistant turns with or without tool calls.

Same fields as chat completion, with one difference: tools is required here, and it can’t be empty. An agentic loop without tools has no results to feed back, so a tool-free task belongs in plain chat completion.

{
  "task_description": "You are a research assistant. Answer the user's question by calling the available tools, reading their results, and replying with a concrete answer grounded in what came back.",
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Look up the current weather for a location",
        "parameters": {
          "type": "object",
          "properties": {
            "location": {
              "type": "string",
              "description": "City to look up the weather for"
            }
          },
          "required": ["location"]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "search",
        "description": "Search for information",
        "parameters": {
          "type": "object",
          "properties": {
            "query": {
              "type": "string",
              "description": "Search query"
            }
          },
          "required": ["query"]
        }
      }
    }
  ]
}

Each row is one whole conversation in a single messages array. Three roles:

  • user turns carry the request, in content.
  • assistant turns carry content, tool_calls, or both, exactly as in chat completion.
  • tool turns carry the result of a call: {"role": "tool", "tool_call_id": ..., "content": ...}. They may only follow an assistant turn that actually made tool calls, and after one, the conversation can continue with another tool result, an assistant turn, or the next user turn.

A conversation starts with a user message and ends with an assistant message — typically the final text answer after the tool results. No system messages; the task_description plays that part.

The tool_call_id on a result points at the id of the call it answers. A single call and a single result are matched automatically; with parallel calls the ids must match. Results are optional per call: a turn can make two calls and carry one result. The examples below set ids on every call.

Written out, one conversation looks like this:

[
  {"role": "user", "content": "What's the weather in Paris?"},
  {"role": "assistant", "content": "Let me check that for you.", "tool_calls": [
    {"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Paris"}}}
  ]},
  {"role": "tool", "tool_call_id": "call_1", "content": "18C, light rain"},
  {"role": "assistant", "content": "It's 18C in Paris with light rain."}
]

In train.jsonl that whole array goes on one line. Cover the loop shapes you need: one call then an answer, chained calls where the second depends on the first’s result, parallel calls, and follow-up user turns after an answered round. Aim for 20+ conversations.

{"messages": [{"role": "user", "content": "What's the weather in Paris?"}, {"role": "assistant", "content": "Let me check that for you.", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Paris"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "18C, light rain"}, {"role": "assistant", "content": "It's 18C in Paris with light rain."}]}
{"messages": [{"role": "user", "content": "Look up the tallest building, then tell me the weather there."}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "search", "arguments": {"query": "tallest building in the world"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "Burj Khalifa, Dubai"}, {"role": "assistant", "tool_calls": [{"id": "call_2", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Dubai"}}}]}, {"role": "tool", "tool_call_id": "call_2", "content": "34C, sunny"}, {"role": "assistant", "content": "It's the Burj Khalifa in Dubai, where it's 34C and sunny."}]}
{"messages": [{"role": "user", "content": "Compare the weather in Rome and Madrid."}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Rome"}}}, {"id": "call_2", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Madrid"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "22C, clear"}, {"role": "tool", "tool_call_id": "call_2", "content": "27C, sunny"}, {"role": "assistant", "content": "Madrid is warmer at 27C and sunny; Rome is 22C and clear."}]}

A conversation can be split into one example per assistant turn, each predicted from everything before it, tool results included. That matches how the deployed model runs: your code executes the calls, appends the results, and calls the model again (Conversation expansion).

As in chat completion, parallel calls are valid data, but generation defaults to one call per turn; raise synthgen.max_tool_calls_per_turn if your task needs more.

Same format as train.jsonl, held out for evaluation. Cover each tool, a multi-step chain, and a final-answer turn. The main score is llm-as-a-judge-reference-free, which judges the text the model writes after the tool results, alongside the tool-call metrics against the reference (Metrics).

{"messages": [{"role": "user", "content": "Is it warm enough for the beach in Barcelona?"}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Barcelona"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "29C, sunny"}, {"role": "assistant", "content": "Yes: 29C and sunny in Barcelona, a good beach day."}]}
{"messages": [{"role": "user", "content": "Where are the next Olympics, and what's the weather there?"}, {"role": "assistant", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "search", "arguments": {"query": "next Olympics host city"}}}]}, {"role": "tool", "tool_call_id": "call_1", "content": "Brisbane, Australia (2032)"}, {"role": "assistant", "tool_calls": [{"id": "call_2", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Brisbane"}}}]}, {"role": "tool", "tool_call_id": "call_2", "content": "24C, partly cloudy"}, {"role": "assistant", "content": "They are in Brisbane, where it's 24C and partly cloudy."}]}

Parallel calls are scored in order, so put the reference calls in the order you want them made.

The task type, the two models and the judge:

base:
  task: chat-completion-agentic
  student_model_name: Qwen3.5-4B
  teacher_model_name: zai.glm-5.3-low-thinking
evaluation:
  llm_as_a_judge_model_name: zai.glm-5.3-low-thinking

Because tools are always declared here, both models have to be tool-calling capable (Supported models).

synthgen.max_tool_calls_per_turn works as in chat completion: unset means one call per generated turn, an integer or unlimited allows more.

To spread the generated data across your tools, add a mutator with one value per tool, as on Chat completion.

Every other field has a default. See Config file for the full table.

Put the files in one directory, here input-dir, and create the seed dataset from it. This validates your files at the same time:

distil seed-dataset create --data input-dir

The command prints Upload successful. Seed dataset ID: <seed-dataset-id>. Run teacher evaluation on that ID next.