Skip to content

Classification data preparation

The example on this page sorts customer service requests into categories for an imaginary bank.

The shared rules for the input directory, row shapes and validation are in Overview. This page covers what’s specific to classification.

What you expect the model to do, in the words you’d use to prompt an LLM. Classification needs two fields:

  • task_description. The task itself.
  • classes_description. Each class name mapped to a description of when it applies. The teacher reads these when generating examples, so a vague description produces vague examples of that class.

llm_as_a_judge_instructions isn’t valid here. Classification is scored on label accuracy, so there’s no judge to instruct.

{
  "task_description": "Classify the bank customer service requests into one of the provided classes",
  "classes_description": {
    "balance_not_updated_after_bank_transfer": "Requests about a completed bank transfer not yet reflected in the account balance. The funds have been debited but not credited, indicating a delay in processing the outgoing transfer.",
    "balance_not_updated_after_cheque_or_cash_deposit": "Requests regarding a recent cheque or cash deposit not showing up in the available account balance. The customer's ledger balance does not reflect the deposit after some time has passed.",
    "card_payment_fee_charged": "Requests questioning an unexpected or additional fee charged for making a payment or purchase with a debit or credit card. The customer seeks clarification on the reason for the fee.",
    "cash_withdrawal_charge": "Requests related to being charged a fee for withdrawing cash from an ATM. The customer wants to know the reason, the exact fee amount, and if it can be waived.",
    "declined_cash_withdrawal": "Requests about attempting to withdraw cash from an ATM but having the transaction declined. The customer has tried multiple ATMs but still faces the same issue with their card.",
    "direct_debit_payment_not_recognised": "Requests regarding an unauthorized direct debit payment charged to the account. The customer claims they did not set it up and wants the bank to investigate its validity."
  }
}

The examples you’re teaching from. Each row is a messages conversation: a user turn holding the input text, and an assistant turn holding the class label.

Aim for 20+ diverse rows, with at least two per class. The teacher generates thousands more from these.

{"messages": [{"role": "user", "content": "Why is there a fee for getting cash?"}, {"role": "assistant", "content": "cash_withdrawal_charge"}]}
{"messages": [{"role": "user", "content": "I was declined when I tried to take out cash!"}, {"role": "assistant", "content": "declined_cash_withdrawal"}]}
{"messages": [{"role": "user", "content": "I deposited some money, but the balance has not changed."}, {"role": "assistant", "content": "balance_not_updated_after_cheque_or_cash_deposit"}]}
{"messages": [{"role": "user", "content": "It has been a couple of hours but I do not see my balance updated, can you help?"}, {"role": "assistant", "content": "balance_not_updated_after_bank_transfer"}]}
{"messages": [{"role": "user", "content": "There is a payment showing on my app that I didn't do. Will you please cancel this payment and refund my money?"}, {"role": "assistant", "content": "direct_debit_payment_not_recognised"}]}
{"messages": [{"role": "user", "content": "How do I know which payments I make will have additional fees? Where can I find this information online?"}, {"role": "assistant", "content": "card_payment_fee_charged"}]}
{"messages": [{"role": "user", "content": "how come i was declined"}, {"role": "assistant", "content": "declined_cash_withdrawal"}]}
{"messages": [{"role": "user", "content": "I deposited a cheque and its been days and I still haven't received the cash!!"}, {"role": "assistant", "content": "balance_not_updated_after_cheque_or_cash_deposit"}]}
{"messages": [{"role": "user", "content": "You promised no fees but now I've got one. What the hell!?"}, {"role": "assistant", "content": "card_payment_fee_charged"}]}
{"messages": [{"role": "user", "content": "Why is there a direct debit to my account? I didn't do that."}, {"role": "assistant", "content": "direct_debit_payment_not_recognised"}]}
{"messages": [{"role": "user", "content": "Do cash withdrawals cost anything?"}, {"role": "assistant", "content": "cash_withdrawal_charge"}]}
{"messages": [{"role": "user", "content": "I sent money from my other bank this morning and my balance here still hasn't moved."}, {"role": "assistant", "content": "balance_not_updated_after_bank_transfer"}]}

Same format as train.jsonl, held out for evaluation. This is what every score you’ll see is measured against: the teacher evaluation, the base student and the tuned student.

Two rules fail validation if you break them:

  • No row may be identical to a training row.
  • Every class in classes_description has to appear here too.

Make it representative of production rather than easy, since a test set of only clear-cut cases gives you a number that won’t survive contact with real traffic.

{"messages": [{"role": "user", "content": "My transfer went out yesterday but the balance still looks wrong."}, {"role": "assistant", "content": "balance_not_updated_after_bank_transfer"}]}
{"messages": [{"role": "user", "content": "Paid in a cheque on Friday, still nothing showing."}, {"role": "assistant", "content": "balance_not_updated_after_cheque_or_cash_deposit"}]}
{"messages": [{"role": "user", "content": "There's an extra charge on top of what I paid at the shop, why?"}, {"role": "assistant", "content": "card_payment_fee_charged"}]}
{"messages": [{"role": "user", "content": "Got charged just for taking my own money out of a machine. Can that be refunded?"}, {"role": "assistant", "content": "cash_withdrawal_charge"}]}
{"messages": [{"role": "user", "content": "Tried three different cash machines and every one refused my card."}, {"role": "assistant", "content": "declined_cash_withdrawal"}]}
{"messages": [{"role": "user", "content": "Something called a direct debit came out today and I never agreed to it."}, {"role": "assistant", "content": "direct_debit_payment_not_recognised"}]}

Unstructured data steers the teacher toward diverse, domain-specific examples. It can be documentation, unlabelled examples, or industry material covering the same ground.

This example uses unlabelled customer requests as raw material for generated examples. Rows carry a single context field.

{"context":"Canceling my order is what I need to do right now."}
{"context":"I swear that there are 2 payments on the app that I didn't make. Could my card be stolen? Please advise what I should do."}
{"context":"To many charges on my card how do I go about fixing that?"}
{"context":"I got less cash than what I specified at the ATM."}
{"context":"Why is my last cheque deposit taking so long?"}
{"context":"A payment shows up on the app that I never made."}
{"context":"My cash withdrawal was short."}
{"context":"I would like to receive a refund for something I bought."}
{"context":"I saw a fee on my receipt from withdrawing money while shopping earlier. Is there supposed to be a fee for this type of transaction?"}
{"context":"I normally don't use ATMs, but I was in a rush today and had to withdraw some cash. The ATM gave me the wrong amount of money and now my app is not showing the same amount as it should. What do I do?"}

The task type, plus the two models:

base:
  task: classification
  student_model_name: Qwen3-0.6B
  teacher_model_name: openai.gpt-oss-120b

Every other field has a default. See Config file for the full table and Supported models for the values you can use.

Create the seed dataset, which validates your files at the same time:

distil seed-dataset create --data ./your-data-dir

Then run teacher evaluation.