Skip to content

Chat completion data preparation

The example on this page is a support assistant for a weather service: it chats with the user in free text and calls a tool when one helps answer the request.

The shared rules for the input directory, row shapes and validation are in Overview. This page covers what’s specific to chat completion.

Chat completion is the conversational task without tool results: the model sees user and assistant turns only. If your conversations feed tool outputs back to the model, use Agentic chat completion instead.

What you expect the model to do, in the words you’d use to prompt an LLM:

  • task_description. The task itself. There is no system-prompt field anywhere in the input: this text becomes the system prompt the model is trained and evaluated with, so write it with that care.
  • tools. Optional. Every tool the model can call, in OpenAI function-calling format, with unique names. Leave it out for a pure chat model with no tools. If your data contains a tool call, the tool has to be declared here, and every call validates against these schemas.
  • llm_as_a_judge_instructions. Optional. Predictions are free text scored by an LLM judge, so this field does real work here.
{
  "task_description": "You are a friendly assistant for a weather service. Chat with the user, and call a tool when it helps answer the request. Answer briefly and concretely, and stay on the topics of weather and travel planning.",
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Look up the current weather for a location",
        "parameters": {
          "type": "object",
          "properties": {
            "location": {
              "type": "string",
              "description": "City to look up the weather for"
            }
          },
          "required": ["location"]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "search",
        "description": "Search for information",
        "parameters": {
          "type": "object",
          "properties": {
            "query": {
              "type": "string",
              "description": "Search query"
            }
          },
          "required": ["query"]
        }
      }
    }
  ]
}

Each row is one whole conversation in a single messages array. Two roles only:

  • user turns carry the request, in content.
  • assistant turns carry content, tool_calls, or both. A turn with neither fails validation.

A conversation starts with a user message, ends with an assistant message, and the two roles alternate strictly. No system messages (the task description plays that part), and no tool messages; tool results are what separates this task from Agentic chat completion.

arguments is a real JSON object, not a JSON-encoded string. That’s the HuggingFace format, and it differs from OpenAI’s chat completions, where arguments is a string. The schemas in job_description.json still use the OpenAI shape, and only the emitted call changes.

Written out, one conversation looks like this:

[
  {"role": "user", "content": "What's the weather in Paris?"},
  {"role": "assistant", "content": "Let me check that for you.", "tool_calls": [
    {"type": "function", "function": {"name": "get_weather", "arguments": {"location": "Paris"}}}
  ]},
  {"role": "user", "content": "And is there anything worth seeing there this weekend?"},
  {"role": "assistant", "tool_calls": [
    {"type": "function", "function": {"name": "search", "arguments": {"query": "Paris events this weekend"}}}
  ]}
]

In train.jsonl that whole array goes on one line. Mix the shapes your production conversations take: plain chat turns, turns that only call a tool, and turns that talk and call at once. Aim for 20+ conversations.

{"messages": [{"role": "user", "content": "What's the weather in Paris?"}, {"role": "assistant", "content": "Let me check that for you.", "tool_calls": [{"type": "function", "function": {"name": "get_weather", "arguments": {"location": "Paris"}}}]}]}
{"messages": [{"role": "user", "content": "Hi! Can you help me plan a trip?"}, {"role": "assistant", "content": "Of course. Where are you headed, and when?"}, {"role": "user", "content": "Lisbon, next week."}, {"role": "assistant", "content": "Nice choice. Let me look up the weather so we can plan around it.", "tool_calls": [{"type": "function", "function": {"name": "get_weather", "arguments": {"location": "Lisbon"}}}]}]}
{"messages": [{"role": "user", "content": "Thanks, that's all I needed."}, {"role": "assistant", "content": "Happy to help. Ask me any time."}]}
{"messages": [{"role": "user", "content": "Compare the weather in Rome and Madrid for me."}, {"role": "assistant", "content": "I'll check both.", "tool_calls": [{"id": "call_1", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Rome"}}}, {"id": "call_2", "type": "function", "function": {"name": "get_weather", "arguments": {"location": "Madrid"}}}]}]}

The last row makes two calls in one turn. Parallel calls are valid data, but by default the teacher generates at most one call per assistant turn; raise synthgen.max_tool_calls_per_turn if your task needs more (see Config below).

Same format as train.jsonl, held out for evaluation. This is what every score you’ll see is measured against: the teacher evaluation, the base student and the tuned student.

Cover the shapes that matter to you: turns that should chat, turns that should call a tool, and follow-ups that depend on earlier turns.

{"messages": [{"role": "user", "content": "Do I need an umbrella in London today?"}, {"role": "assistant", "content": "Let me look at the current weather.", "tool_calls": [{"type": "function", "function": {"name": "get_weather", "arguments": {"location": "London"}}}]}]}
{"messages": [{"role": "user", "content": "What can you do?"}, {"role": "assistant", "content": "I can check the current weather for any city and search for travel information. What do you need?"}]}
{"messages": [{"role": "user", "content": "Find me something to do in Berlin."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "search", "arguments": {"query": "things to do in Berlin"}}}]}]}

Evaluation expands each conversation into one line per assistant turn, each predicted from the turns before it, so the test-set size reported back to you won’t match the number of rows you uploaded. Data that’s already in this split form is detected and passed through unchanged.

Parallel calls are scored in order: a turn with the right calls in the wrong order counts as a miss, so put the reference calls in the order you want them made.

The task type, plus the two models:

base:
  task: chat-completion
  student_model_name: Qwen3-1.7B
  teacher_model_name: openai.gpt-oss-120b

If the job description declares tools, both models have to be tool-calling capable: students are limited to the Qwen3, Qwen3.5, Llama 3-family, LFM2/LFM2.5, FunctionGemma and Gemma 4 families, and teachers to those marked in the tool-calling column of Supported models. With no tools declared, any supported model works.

One parameter is specific to this task family:

synthgen:
  max_tool_calls_per_turn: unlimited

It caps how many tool calls a generated assistant turn may contain. Unset, the limit is 1 when tools are declared and 0 when they aren’t. Set an integer or unlimited to allow parallel calls, or 0 to forbid calls outright. Declaring no tools while setting it above zero fails validation, as does declaring tools while setting it to 0.

Every other field has a default. See Config file for the full table and Supported models for the values you can use.

Create the seed dataset, which validates your files at the same time:

distil seed-dataset create --data ./your-data-dir

Then run teacher evaluation.