Skip to content

Multi-turn tool calling data preparation

The example on this page navigates a file system: given a conversation of user requests and assistant tool calls, the model produces the next call.

The shared rules for the input directory, row shapes and validation are in Overview. This page covers what’s specific to multi-turn tool calling.

What you expect the model to do, in the words you’d use to prompt an LLM. Tool calling needs two fields:

  • task_description. The task itself.
  • tools. Every tool the model can call, in OpenAI function-calling format, with unique names. Every call in your training and test data validates against these schemas, so a call to a tool you didn’t declare fails the job.

llm_as_a_judge_instructions isn’t valid here. Tool calls are scored against the reference call, so there’s no judge to instruct.

{
  "task_description": "You are an expert agent that navigates and manages files in a file system. Given a command or request from the user, call the appropriate file system function to complete the request.",
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "ls",
        "description": "List the contents of the current directory.",
        "parameters": {
          "type": "object",
          "properties": {
            "a": {
              "type": "boolean",
              "description": "Show hidden files and directories. Defaults to False.",
              "default": false
            }
          },
          "required": []
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "cd",
        "description": "Change the current working directory to the specified folder.",
        "parameters": {
          "type": "object",
          "properties": {
            "folder": {
              "type": "string",
              "description": "The folder of the directory to change to."
            }
          },
          "required": ["folder"]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "cat",
        "description": "Display the contents of a file from the current directory.",
        "parameters": {
          "type": "object",
          "properties": {
            "file_name": {
              "type": "string",
              "description": "The name of the file to display."
            }
          },
          "required": ["file_name"]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "mkdir",
        "description": "Create a new directory in the current directory.",
        "parameters": {
          "type": "object",
          "properties": {
            "dir_name": {
              "type": "string",
              "description": "The name of the new directory to create."
            }
          },
          "required": ["dir_name"]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "touch",
        "description": "Create a new file in the current directory.",
        "parameters": {
          "type": "object",
          "properties": {
            "file_name": {
              "type": "string",
              "description": "The name of the new file to create."
            }
          },
          "required": ["file_name"]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "rm",
        "description": "Remove a file or directory.",
        "parameters": {
          "type": "object",
          "properties": {
            "file_name": {
              "type": "string",
              "description": "The name of the file or directory to remove."
            }
          },
          "required": ["file_name"]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "cp",
        "description": "Copy a file or directory from one location to another.",
        "parameters": {
          "type": "object",
          "properties": {
            "source": {
              "type": "string",
              "description": "The name of the file or directory to copy."
            },
            "destination": {
              "type": "string",
              "description": "The destination name to copy to."
            }
          },
          "required": ["source", "destination"]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "mv",
        "description": "Move or rename a file or directory.",
        "parameters": {
          "type": "object",
          "properties": {
            "source": {
              "type": "string",
              "description": "Source name of the file or directory to move."
            },
            "destination": {
              "type": "string",
              "description": "The destination name to move to."
            }
          },
          "required": ["source", "destination"]
        }
      }
    }
  ]
}

Each row is one whole conversation in a single messages array. The earlier turns are the history that gives the model context, and the last assistant turn is the call it’s trained to produce.

  • user turns carry the request, in content.
  • assistant turns carry exactly one call in tool_calls and omit content. An empty string is also accepted. More than one call per turn isn’t supported.
  • tool turns are optional. They carry the result of the call before them, and can be followed by either the next user turn or the next assistant call.

arguments is a real JSON object, not a JSON-encoded string. That’s the HuggingFace format, and it differs from OpenAI’s chat completions, where arguments is a string. The schemas in job_description.json still use the OpenAI shape, and only the emitted call changes.

Written out, one conversation looks like this:

[
  {"role": "user", "content": "Please list all the files in my current directory."},
  {"role": "assistant", "tool_calls": [
    {"type": "function", "function": {"name": "ls", "arguments": {}}}
  ]},
  {"role": "user", "content": "Show me what's inside config.txt."},
  {"role": "assistant", "tool_calls": [
    {"type": "function", "function": {"name": "cat", "arguments": {"file_name": "config.txt"}}}
  ]}
]

In train.jsonl that whole array goes on one line. Aim for 20+ conversations covering every tool and the ways they get chained.

{"messages": [{"role": "user", "content": "Please list all the files in my current directory."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "ls", "arguments": {}}}]}, {"role": "user", "content": "Navigate to the backup directory."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "cd", "arguments": {"folder": "backup"}}}]}, {"role": "user", "content": "Create a new folder called archive_2024."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "mkdir", "arguments": {"dir_name": "archive_2024"}}}]}, {"role": "user", "content": "Show me what's inside config.txt."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "cat", "arguments": {"file_name": "config.txt"}}}]}]}
{"messages": [{"role": "user", "content": "What files are in here?"}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "ls", "arguments": {}}}]}, {"role": "user", "content": "Go to the downloads folder."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "cd", "arguments": {"folder": "downloads"}}}]}, {"role": "user", "content": "Copy invoice_march.pdf to the accounting folder."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "cp", "arguments": {"source": "invoice_march.pdf", "destination": "accounting"}}}]}]}
{"messages": [{"role": "user", "content": "Show me all files here including hidden ones."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "ls", "arguments": {"a": true}}}]}, {"role": "user", "content": "Change directory to research."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "cd", "arguments": {"folder": "research"}}}]}, {"role": "user", "content": "Create a file called experiment_log.txt."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "touch", "arguments": {"file_name": "experiment_log.txt"}}}]}]}
{"messages": [{"role": "user", "content": "Display the contents of my current directory."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "ls", "arguments": {}}}]}, {"role": "user", "content": "Move to the utils directory."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "cd", "arguments": {"folder": "utils"}}}]}, {"role": "user", "content": "Remove helper.js."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "rm", "arguments": {"file_name": "helper.js"}}}]}]}
{"messages": [{"role": "user", "content": "List files in the current location."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "ls", "arguments": {}}}]}, {"role": "user", "content": "Navigate to the media folder."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "cd", "arguments": {"folder": "media"}}}]}, {"role": "user", "content": "Move photo_001.jpg to the gallery directory."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "mv", "arguments": {"source": "photo_001.jpg", "destination": "gallery"}}}]}]}

A conversation can also carry the output of a call, so the next request can depend on what came back:

{"messages": [{"role": "user", "content": "Show me what's inside notes.txt."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "cat", "arguments": {"file_name": "notes.txt"}}}]}, {"role": "tool", "content": "Buy milk. Call the plumber."}, {"role": "user", "content": "Now delete the file."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "rm", "arguments": {"file_name": "notes.txt"}}}]}]}

Same format as train.jsonl, held out for evaluation. This is what every score you’ll see is measured against: the teacher evaluation, the base student and the tuned student.

Three rules fail validation if you break them:

  • No row may be identical to a training row.
  • Every call has to validate against your tools schemas.
  • Each conversation starts with a user message, ends with an assistant tool call, and alternates roles validly.

Cover each tool at least once. Follow-ups that depend on an earlier turn are the thing worth testing here, since that dependency is what separates this task from single-turn tool calling.

{"messages": [{"role": "user", "content": "What's in the current directory?"}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "ls", "arguments": {}}}]}, {"role": "tool", "content": "reports/  archive/  notes.txt"}, {"role": "user", "content": "Open the notes file."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "cat", "arguments": {"file_name": "notes.txt"}}}]}]}
{"messages": [{"role": "user", "content": "Make a folder called drafts."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "mkdir", "arguments": {"dir_name": "drafts"}}}]}, {"role": "tool", "content": "Created drafts"}, {"role": "user", "content": "Now move notes.txt into it."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "mv", "arguments": {"source": "notes.txt", "destination": "drafts"}}}]}]}
{"messages": [{"role": "user", "content": "Show me everything, hidden files included."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "ls", "arguments": {"a": true}}}]}, {"role": "user", "content": "Switch to the .cache folder."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "cd", "arguments": {"folder": ".cache"}}}]}, {"role": "user", "content": "Delete stale.bin from here."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "rm", "arguments": {"file_name": "stale.bin"}}}]}]}
{"messages": [{"role": "user", "content": "Go into the templates directory."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "cd", "arguments": {"folder": "templates"}}}]}, {"role": "user", "content": "Start a new file named header.html."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "touch", "arguments": {"file_name": "header.html"}}}]}, {"role": "user", "content": "Duplicate it into the backup folder."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "cp", "arguments": {"source": "header.html", "destination": "backup"}}}]}]}

Evaluation expands each conversation into one line per tool call, so the test-set size reported back to you won’t match the number of rows you uploaded.

Unstructured data steers the teacher toward diverse, domain-specific examples. Here that means material describing the environment the tools act on, so generated conversations use realistic paths and file names. Rows carry a single context field.

{"context": "Project layout: src/ holds the application code, tests/ mirrors it one file per module, docs/ holds the published guides, and build/ is generated output that can be removed at any time."}
{"context": "Backup convention: nightly archives land in backup/ named archive_YYYY-MM-DD, and anything older than 30 days is moved to cold-storage/ rather than deleted."}
{"context": "Hidden files in a home directory usually include .bashrc, .gitconfig and .cache. The .cache directory can grow to several gigabytes and is safe to empty."}
{"context": "Media assets are stored under media/ as photo_NNN.jpg, and published copies are moved into gallery/ once they have been reviewed."}

The task type, plus the two models:

base:
  task: multi-turn-tool-calling-closed-book
  student_model_name: Qwen3-1.7B
  teacher_model_name: openai.gpt-oss-120b

Tool calling restricts both. Students are limited to the Qwen3, Qwen3.5, Llama 3-family, LFM2/LFM2.5, FunctionGemma and Gemma 4 families, and teachers to those marked in the tool-calling column of Supported models. The default teacher, openai.gpt-oss-120b, works here.

Every other field has a default. See Config file for the full table and Supported models for the values you can use.

Create the seed dataset, which validates your files at the same time:

distil seed-dataset create --data ./your-data-dir

Then run teacher evaluation.