What Is Question Answering as a Training Task?
Question answering is the distil labs task type for any input-to-text problem where the output is a targeted answer rather than a label or a function call. You declare task: question-answering, and the model learns to locate or generate the specific value asked for instead of paraphrasing the whole input.
How does the question answering task work?
Each training example is a two-turn conversation: a user turn holding the input and the question, an assistant turn holding the expected answer. Everything the model needs sits in that user turn — plain question-answering has no separate context field and no retrieval step.
The question answering data preparation guide uses invoice extraction as the worked example, and the shape is instructive:
{"messages": [{"role": "user", "content": "Invoice #1234 from Acme Corp dated 2024-01-15. Total: $540. What is the total amount?"}, {"role": "assistant", "content": "$540"}]}
The answer is $540, not a sentence about the invoice. That terseness is the point: you are training the output contract as much as the extraction. If your downstream code expects a bare value, every seed answer should be a bare value.
This task type covers far more than literal questions. The docs frame it as the fit for transforming text “by addressing implicit questions about its content”, which is where summarisation, rewriting, and structured extraction land — the name describes the input-output shape, not the subject.
What goes in the job description?
Two required fields and one optional one, and the optional one matters more than its name suggests.
| Field | Required | What it does |
|---|---|---|
task_description |
yes | What the model should do and how to format output, including the fallback for missing values |
input_description |
yes | What the input looks like, so the teacher can fabricate realistic new inputs |
llm_as_a_judge_instructions |
no | The rubric the judge applies when scoring predictions |
task_description is where you pin the output contract: “Return only the specific value asked for, without additional explanation. If the information is not found, respond with ‘Not found’.” Without a stated fallback, the model will invent one, and inventing is the failure mode you are training away.
llm_as_a_judge_instructions is worth writing even though it is optional, because it decides what counts as correct. A rubric that accepts semantic equivalence and a rubric that demands the exact reference string will grade the same model very differently.
Which question answering variant do you actually need?
There are three, and plain question-answering is the one where the answer is already in the input you send.
| Variant | base.task |
Where the knowledge comes from |
|---|---|---|
| Question answering | question-answering |
The input text itself |
| Open-book QA (RAG) | question-answering-open-book |
A context passage supplied alongside the question |
| Closed-book QA | question-answering-closed-book |
The model’s weights, learned during training |
The distinction between the first two is structural, not philosophical: open-book examples carry a sibling context field in each JSONL line, so the platform knows which text the answer must be grounded in and can generate distractor blocks against it. If you paste the document into the user turn yourself, you are doing plain question answering. Open-book vs closed-book QA covers the second choice in detail.
How is a question answering model scored?
With three metrics that disagree on purpose. The metrics reference recommends LLM-as-a-judge as the headline number.
- LLM-as-a-judge — a large model grades the prediction against the reference and the rubric. Best when many phrasings are correct; the research behind the approach is surveyed in this arXiv paper.
- Exact match — 1 if the strings match, 0 otherwise. Honest and harsh, useful when there is one correct phrasing.
- ROUGE-L — longest-common-subsequence overlap. Rewards reuse of reference wording and favours longer answers.
Read them together. Low exact match with a high judge score means the answers are right but paraphrased, which is a formatting problem rather than an accuracy problem. All three low means the task itself is under-specified — revisit task_description before you revisit the model.
Related terms
Teacher evaluation runs a large model over your test set before training to check the task is solvable at all; see teacher evaluation. Unstructured data is in-domain text that steers synthetic generation toward realistic inputs. Mutation topics (synthgen.mutation_topics in the config file) push the teacher to generate different kinds of question rather than a thousand variations of the easiest one. For related reading, what is a small language model covers the students, and when does distillation fail covers what goes wrong when a QA task is too broad.