Skip to content

Question answering data preparation

The example on this page extracts information from invoices: the input carries the invoice text and the question, and the model returns the value asked for.

The shared rules for the input directory, row shapes and validation are in Overview. This page covers what’s specific to question answering.

What you expect the model to do, in the words you’d use to prompt an LLM.

  • task_description. What the model should do and how it should format its output.
  • synthetic_data_generation_instructions. Optional. What the generated inputs should look like: their format, the fields they carry, how much they vary. It’s read during generation and nowhere else, so it steers the inputs without redefining the task.
  • llm_as_a_judge_instructions. Optional, but it does real work here: this task is scored by an LLM judge, so these instructions decide what counts as correct. State what has to match and what to ignore.
{
    "task_description": "Extract the requested information from the provided invoice text. Return only the specific value asked for, without additional explanation. If the information is not found, respond with 'Not found'.",
    "synthetic_data_generation_instructions": "Each input is one invoice text followed by a single question about that invoice. The invoice text carries an invoice number, a date, a vendor name, line items, a subtotal, tax and a total amount. Vary the vendors, the number of line items and the field the question asks for, and include invoices where the requested value is absent.",
    "llm_as_a_judge_instructions": "Compare the predicted answer to the reference answer for the given question. Output 'good' if the prediction matches the reference value or is semantically equivalent, otherwise output 'bad'."
}

Each row is a messages conversation: a user turn holding the input text and the question, and an assistant turn holding the expected answer.

Aim for 20+ diverse rows. The teacher generates thousands more from these.

{"messages": [{"role": "user", "content": "Invoice #1234 from Acme Corp dated 2024-01-15. Items: Widget x10 at $50 each. Subtotal: $500. Tax: $40. Total: $540. What is the total amount?"}, {"role": "assistant", "content": "$540"}]}
{"messages": [{"role": "user", "content": "Invoice #1234 from Acme Corp dated 2024-01-15. Items: Widget x10 at $50 each. Subtotal: $500. Tax: $40. Total: $540. What is the invoice number?"}, {"role": "assistant", "content": "1234"}]}
{"messages": [{"role": "user", "content": "Invoice #1234 from Acme Corp dated 2024-01-15. Items: Widget x10 at $50 each. Subtotal: $500. Tax: $40. Total: $540. Who is the vendor?"}, {"role": "assistant", "content": "Acme Corp"}]}
{"messages": [{"role": "user", "content": "Invoice #5678 from Global Services dated 2024-02-20. Items: Consulting 8hrs at $150/hr. Subtotal: $1200. Tax: $0. Total: $1200. What is the invoice date?"}, {"role": "assistant", "content": "2024-02-20"}]}
{"messages": [{"role": "user", "content": "Invoice #5678 from Global Services dated 2024-02-20. Items: Consulting 8hrs at $150/hr. Subtotal: $1200. Tax: $0. Total: $1200. What is the tax amount?"}, {"role": "assistant", "content": "$0"}]}

Same format as train.jsonl, held out for evaluation. This is what every score you’ll see is measured against: the teacher evaluation, the base student and the tuned student.

No row may be identical to a training row, or validation fails. Use different invoices rather than different questions about the same ones, so the test set measures generalisation instead of recall.

Cover the awkward cases too. The example below includes a value that isn’t present, which is what tells you whether the model follows the “respond with ‘Not found’” rule under pressure.

{"messages": [{"role": "user", "content": "Invoice #4471 from Northwind Traders dated 2024-05-02. Items: Freight x1 at $220. Subtotal: $220. Tax: $17.60. Total: $237.60. What is the total amount?"}, {"role": "assistant", "content": "$237.60"}]}
{"messages": [{"role": "user", "content": "Invoice #4471 from Northwind Traders dated 2024-05-02. Items: Freight x1 at $220. Subtotal: $220. Tax: $17.60. Total: $237.60. Who is the vendor?"}, {"role": "assistant", "content": "Northwind Traders"}]}
{"messages": [{"role": "user", "content": "Invoice #8802 from Bright Labs dated 2024-06-19. Items: Assay kit x3 at $410 each. Subtotal: $1230. Tax: $98.40. Total: $1328.40. What is the invoice date?"}, {"role": "assistant", "content": "2024-06-19"}]}
{"messages": [{"role": "user", "content": "Invoice #8802 from Bright Labs dated 2024-06-19. Items: Assay kit x3 at $410 each. Subtotal: $1230. Tax: $98.40. Total: $1328.40. What is the purchase order number?"}, {"role": "assistant", "content": "Not found"}]}
{"messages": [{"role": "user", "content": "Invoice #1590 from Harbour Print dated 2024-07-08. Items: Poster run x500 at $1.20 each. Subtotal: $600. Tax: $0. Total: $600. What is the tax amount?"}, {"role": "assistant", "content": "$0"}]}

Unstructured data steers the teacher toward diverse, domain-specific examples. For question answering, supply realistic samples of the inputs your model will meet in production. Rows carry a single context field.

{"context": "Invoice #9012 from Tech Solutions Inc dated 2024-03-10. Items: Software License x1 at $299. Subtotal: $299. Tax: $24. Total: $323."}
{"context": "Invoice #3456 from Office Supplies Co dated 2024-03-15. Items: Paper 10 reams at $8 each, Pens box x5 at $12 each. Subtotal: $140. Tax: $11. Total: $151."}
{"context": "Invoice #7890 from Cloud Hosting Ltd dated 2024-04-01. Items: Monthly hosting at $99, Domain renewal at $15. Subtotal: $114. Tax: $0. Total: $114."}
{"context": "Invoice #2468 from Marketing Agency dated 2024-04-12. Items: Campaign management 20hrs at $75/hr. Subtotal: $1500. Tax: $120. Total: $1620."}

The task type, plus the two models:

base:
  task: question-answering
  student_model_name: Qwen3-0.6B
  teacher_model_name: openai.gpt-oss-120b

Every other field has a default. See Config file for the full table and Supported models for the values you can use.

Create the seed dataset, which validates your files at the same time:

distil seed-dataset create --data ./your-data-dir

Then run teacher evaluation.