← All content
GuideClassificationQuestion AnsweringAgentic AI
Autoresearch with distil labs: let your agent iterate on building SLMs for your case

Autoresearch with distil labs: let your agent iterate on building SLMs for your case

Your first fine-tuned model is rarely the one you ship, because training it is the easy part: the time goes into the second and third iterations, where you read through the predictions it got wrong, work out whether the training data, the teacher or the training settings are at fault, change one thing, and run the whole pipeline again. Each round is up to an hour of jobs and a lot of careful reading in between, and most of that reading follows the same pattern every time.

So we handed that loop to coding agents. We gave three agents the distil labs agent skill, one task each and a budget of three training runs, and they iterated on their own, with no human in the loop, until their 0.6B student reached its 120B teacher or beat the production service it was meant to replace. Two of the students reached the teacher (1.00 against a teacher at 1.00, and 1.00 against 0.98-1.00), and the third beat the production service (0.709 against 0.663).

What autoresearch is

distil labs is a platform that fine-tunes task-specific small language models automatically. Most people use it to swap the general-purpose LLM in their agentic system for a smaller, purpose-built one, with the same quality at around 80% lower cost and lower latency.

Autoresearch is your coding agent running the model iterations on its own. It trains a model, reads its scores against the teacher and the untrained base model, inspects the predictions the model got wrong, finds the stage those errors come from, changes one setting there, and trains again. It keeps going until the model reaches a target you agreed on, the budget is spent, or it needs a decision only you can make.

Results

We ran the loop on three of our example tasks, each with a Qwen3-0.6B student, a gpt-oss-120b teacher and a budget of three training runs. Starting from untrained students at 0.58, 0.20 and 0.07, two reached the teacher’s score, on the first and on the second iteration, and the third climbed to 0.71 over three iterations, past the 0.66 of the production service it replaces.

Student score per autoresearch iteration on the three tasks, with the teacher, the base student and the production service as reference lines

Setting up the loop

The loop runs on the distil labs agent skill, which works in Claude Code, Codex, Cursor and the other agents that read the Agent Skills format. You need the distil CLI signed in to your account, the skill installed, and a starting point: traffic collected by an inference endpoint, a file of production traces, or a labeled dataset with a test set.

distil skill install

Then you describe the task and ask for the loop:

Use the distil labs skill to build a model from the traces in ./traces-input, then keep
iterating on it until it is as accurate as you can get it within 3 training runs.

Before the first iteration the agent agrees four things with you: the target, which defaults to the teacher’s score; the budget, in training runs and credits; the judge model and instructions, fixed for every run so the scores stay comparable; and how it measures the noise between runs, which it does by running one evaluation twice. From then on it records every iteration in an iterations.md file in the project directory, which is how you follow along, and how the agent picks up again after its context is reset. Here is the log from one of our runs, shortened:

Iteration Change Primary metric Kept
1 baseline 0.668 yes
2 reuse iteration 1’s data, fix 66 rows that break the override rule, rebalance with a mutator 0.709 yes
3 8 epochs instead of 4 0.693 no

Each iteration goes around the same loop:

  ┌──► read the scores ────► inspect the wrong predictions
  │                                       │
  │                                       ▼
  train again ◄── change one ◄── find the stage the
                  setting there    errors come from

  stops when the target is reached, the budget is spent,
  or it needs a decision from you

How each iteration trains a model

Every iteration runs the same distil labs pipeline. It starts from your seed examples or production traces and a test set, scores the teacher on that test set as the ceiling and the untrained student as the floor, has the teacher generate synthetic training data (256 to 1,000 rows per run in these tests, 10,000 by default), filters it with validators for format, length and duplicates, and fine-tunes the student with LoRA for 4 epochs. A student can match its teacher on a narrow task because the validators strip the teacher’s mistakes and the student spends all of its capacity on one job.

What the agent adds is the reading in between. After each run it downloads the student’s and the teacher’s predictions on the same test rows, groups the student’s errors into named failure modes (“misses one tag on reports with three conditions”, “confuses two neighbouring codes”), and maps each one to the stage that can fix it: wrong test labels go back to the test set, a teacher that gets them wrong too means a different teacher, gaps in the training data go to synthetic data generation, and spread-out errors with a small gap to the teacher go to training. It changes one thing, cheapest stage first, and keeps the change only if the score moves by more than the noise.

Our test: three agents, three tasks

Each agent was a fresh Claude Code session given only the skill and one of our onboarding examples, with the same constraints: Qwen3-0.6B as the student, gpt-oss-120b as the teacher and the judge, at most three full training runs, and no deployment.

  • incident-triage. One raw log line from a monitoring pipeline in, one incident record out: severity, component, whether to page someone, and how long to suppress duplicates. 20 seed examples, 50 test rows.
  • bindery-defect-triage. One inspection line from a bindery’s quality station in, one disposition code out, across 30 defect names of the plant’s own shorthand. 30 seed examples, 50 test rows.
  • trail-report-tagging. One free-text trail report in, the condition tags that apply out, from a vocabulary of 48 codes. It starts from 500 traces logged by a production service that already does this job, and badly.

How the agents progressed

incident-triage: one run

The agent ran the example’s configuration as the baseline and got 1.00 on the 50 test rows, against 0.58 for the untrained student and 1.00 for the teacher in both of its teacher evaluations. It did not take the score on trust: it wrote its own script, checked all 50 predictions against the task’s rule table, found them all correct, and stopped after one of its three runs, about 50 minutes in.

bindery-defect-triage: two runs

This one needed work before the first training run. In its smoke runs of synthetic data generation, the agent re-derived every generated label from the task’s rules and found that the example’s mutators, a grid of station × severity, made the teacher write 8-15% wrong labels; without them, 1 row in 128 was wrong. It removed the grid, tightened the generation instructions (the exact line prefix, the defect names verbatim, a label consistent with the rules), and trained: 0.98, against 0.20 for the base model and a teacher at 0.98 and 1.00 across its two runs.

The one remaining error was a defect name, crossfold, that appeared in 0 of the 542 training rows. For the second iteration the agent reused those 542 rows as the seed and generated 256 more aimed at the defect names, which reached 1.00. That is one row better on 50 rows, within the 0.02 noise it had measured from the teacher’s two runs, and since 1.00 can’t be improved it stopped after two runs, about 107 minutes in.

trail-report-tagging: three runs

Starting from the traces, the agent built a 199-row test set from traces, which also scored the production service that wrote them at 0.663, and processed the rest into a seed dataset. The baseline student got 0.668, against 0.065 for the base model and a teacher at 0.940 and 0.920.

The teacher was right on 55 of the student’s 77 wrong rows, so the agent looked at what the student had learned from: 66 training rows broke the task’s override rule (an impassable trail suppresses the tags for the conditions it causes), and one tag sat on 44% of the training rows against 22% of the test rows. The second iteration reused that training data with the broken rows fixed, added 640 rows weighted toward reports with several conditions, and reached 0.709, up 0.041, with invented tag codes down from 4 to 0. A third iteration with 8 epochs instead of 4 scored 0.693, within the 0.02 noise, so the agent reverted it, about 171 minutes in.

Detailed results

Task Base student Teacher Best student Production service Training runs Agent time
incident-triage 0.58 1.00 ± 0.00 1.00 - 1 of 3 ~50 min
bindery-defect-triage 0.20 0.99 ± 0.01 1.00 - 2 of 3 ~107 min
trail-report-tagging 0.065 0.93 ± 0.01 0.709 0.663 3 of 3 ~171 min

Scores are llm-as-a-judge-reference-free (incident-triage, trail-report-tagging) and accuracy (bindery-defect-triage), on 50, 50 and 199 test rows. The teacher’s ± is the spread between its two runs; each student score is a single run, so a difference within that spread (one row on the 50-row sets) is noise.

On two tasks a Qwen3-0.6B student, 200x smaller than its gpt-oss-120b teacher, reached the teacher’s score. On the third it beat the production service by 0.046 while closing 74% of the gap between the base student and the teacher, so it is the better model to deploy but not yet at the teacher. At these data sizes a training run took about 25-30 minutes and synthetic data generation 3-8 minutes, so each full loop, from the first smoke run to the agent’s final report, finished in 50 to 171 minutes without anyone reading predictions by hand.

Try it on your task

Install the skill with distil skill install, point your coding agent at your production traces or a handful of labeled examples, and give it a budget of training runs. The autoresearch guide has the setup in full.


distil labs trains task-specific small language models that are cheaper and faster per request than general-purpose LLMs, with equal-or-better accuracy on bounded tasks. Docs · GitHub · Slack