Skip to content

Autoresearch

Autoresearch is your coding agent running model iterations on its own: it trains a model, inspects the predictions it got wrong, changes one setting in the stage those errors come from, and trains again, until the model reaches your target or the budget is spent. It runs on the distil labs agent skill.

You need the CLI signed in, the skill installed (distil skill install), a starting point from the build loop, and training credits: every iteration trains at least one model. Then ask your agent, for example:

Use the distil labs skill to build a model from the traces of our inference endpoint
support-yeOdAS, then keep iterating until it is as accurate as you can get it within our credits.

Before it starts, the agent agrees with you on:

  • the primary metric: the default for the task unless you name another;
  • a target on that metric: the teacher’s score, unless you set another one;
  • a budget: credits per route and a maximum number of iterations;
  • the judge, fixed for every run so scores stay comparable.

Each iteration, the agent reads the scores, groups the wrong predictions into failure modes, changes one setting in the stage they come from, and trains again. A change that doesn’t help is reverted.

It records every iteration in iterations.md in the project directory, so you can follow along and the loop survives a context reset:

Iteration Change Primary metric Kept
1 baseline 0.71 yes
2 mutator for refund questions 0.78 yes
3 larger student 0.77 no

It stops when it reaches the target, spends the budget, stops improving, or needs a decision from you, and then reports what each iteration changed and why it stopped.