Skip to content

Legacy data format

If your train and test rows are flat, with question, answer and, for open book QA, context columns, convert them before you upload. distil seed-dataset create accepts only the messages format: one conversation per example, matching the chat format the model is trained and served on. This page maps the flat shape to it, task by task.

  • One messages array per example. The question and answer columns become a conversation: a user turn holding the input, and an assistant turn holding the expected output.
  • Tool calls use the HuggingFace format. The call lives in the assistant turn’s tool_calls array and arguments is a real JSON object. Neither a stringified answer with a parameters key nor OpenAI’s stringified arguments is accepted.

Raw traces are unaffected. What you upload with distil traces upload stays in the OpenAI chat-completions format, including OpenAI-style tool calls where arguments is a string. See Trace inputs.

Question answering, classification, closed book QA

Section titled “Question answering, classification, closed book QA”

Old

{"question": "What is the total amount due?", "answer": "$540"}

New

{"messages": [{"role": "user", "content": "What is the total amount due?"}, {"role": "assistant", "content": "$540"}]}

Old

{"question": "How many students enrolled?", "context": "The university enrolled 5,984 students...", "answer": "5,984"}

New

{"messages": [{"role": "user", "content": "How many students enrolled?"}, {"role": "assistant", "content": "5,984"}], "context": "The university enrolled 5,984 students..."}

The answer string becomes an assistant tool_calls array. arguments is a JSON object (HuggingFace format) rather than a stringified blob, and there’s no parameters key.

Old

{"question": "What's the weather in New York?", "answer": "{\"name\":\"get_weather\",\"parameters\":{\"location\":\"New York, NY\"}}"}

New

{"messages": [{"role": "user", "content": "What's the weather in New York?"}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "get_weather", "arguments": {"location": "New York, NY"}}}]}]}

An assistant turn that makes a tool call omits content. An empty string "" is also accepted.

A stringified JSON array in question, plus a separate answer tool call, become a single messages array. The target tool call is the final assistant turn.

Old

{"question": "[{\"role\": \"user\", \"content\": \"List files here.\"}, {\"role\": \"assistant\", \"content\": \"\", \"tool_calls\": [{\"type\": \"function\", \"function\": {\"name\": \"ls\", \"arguments\": {}}}]}, {\"role\": \"user\", \"content\": \"Show me config.txt.\"}]", "answer": "{\"name\": \"cat\", \"parameters\": {\"file_name\": \"config.txt\"}}"}

New

{"messages": [{"role": "user", "content": "List files here."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "ls", "arguments": {}}}]}, {"role": "user", "content": "Show me config.txt."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "cat", "arguments": {"file_name": "config.txt"}}}]}]}

Watch the two halves, because they don’t use the same key. Assistant turns inside the history carry arguments as a JSON object, while the separate answer carries parameters as a string. Both convert to arguments, so the target call ends up the same shape as the calls before it.

Once converted, the files go in an input directory like any other dataset. The shared rules are in Overview, and your task’s page has a worked example.

Create the seed dataset, which validates the converted files at the same time:

distil seed-dataset create --data ./your-data-dir

Then run teacher evaluation.