Changelog

New features and improvements to distil labs, as they ship. Follow along on LinkedIn, X or in our Slack community.

  1. Auto mutators: let the teacher pick the values

    Mutators let you choose which situations your small model practices during training, including the ones that rarely show up in your production traces. Until now you had to list every value a mutator could take yourself, which meant knowing your data well before you started.

    With auto mutators, you describe in one sentence what should vary, for example which part of the product the customer is asking about. The teacher reads your job description and a sample of your seed data, then proposes the values. They are printed in the run logs, so you can check them and pin down or edit any you disagree with.

    Read the docs →
  2. Reasoning for small models

    Your small model can now reason through a task step by step before it answers. That matters for decisions the input doesn’t state outright, like whether an invoice total adds up, where a model answering in one pass hits a ceiling. During training, a larger teacher writes step-by-step examples, and your small model learns to work through new inputs the same way.

    To try it, set enable_thinking: true before generating training data, and describe the reasoning format and length in synthetic_data_generation_instructions.

    Fine-tuned Qwen3.5-4B got 98 of our 100 test invoices right with reasoning, against 84 when answering directly and 97 for GLM 5.3, the larger model that generated its training data.

    Read the docs →
  3. Inference endpoint: start with the traffic you already have

    Send your product’s traffic through a distil labs inference endpoint and we capture the requests and responses while your existing LLM keeps answering. You no longer need to prepare a trace export before getting started: sign up, create an endpoint and point your application at it.

    Switching over means changing three strings: the base URL, the API key and the model name.

     client = OpenAI(
    +    base_url="https://inference.distillabs.ai/v1",
    -    api_key=os.environ["OPENAI_API_KEY"],
    +    api_key=os.environ["DISTIL_API_KEY"],
     )
    
     client.chat.completions.create(
    -    model="gpt-4.1-mini",
    +    model="support-yeOdAS",
         messages=messages,
     )

    The captured traffic is the starting point for your small model’s training data.

    Read the docs →
  4. Mutators: choose what your model trains on

    Mutators give you control over which situations your synthetic training data covers, including the ones your production traces barely contain.

    The situations where a failure matters most are often rare in production, like a caller switching language halfway through a call or interrupting the agent mid-sentence. Training on your traces, or generating more of what they already contain, leaves your model with few examples of them. With mutators, you describe those situations and how often each should appear, and the teacher generates examples for your small model to learn from.

    Read the docs →
  5. Fully self-serve: a deployed model in about 30 minutes

    Sign up, point the platform at the production traces you already collect, and have a task-specific small model deployed behind an OpenAI-compatible endpoint in about 30 minutes. Most of that time is jobs running on our hardware while you do something else.

    Training runs in six stages, from uploading traces to deployment, and each one returns a score, so you know whether the task is learnable before you spend anything on training. Drive it from the CLI, one command per stage, or install the skill and let your coding agent run it with you.

    No production traces yet? A few dozen examples of the task are enough to try it, and the free tier covers two full training runs.

    Read the docs →
  6. Conversations are now first-class citizens

    You can now fine-tune a small model on the whole conversation, text and tool calls together. We have trained multi-turn tool calling since June 2025, and two new task types cover the rest.

    chat-completion is for an assistant that talks to the user and can fire tool calls in the same turn, so it can say “let me check that for you” and issue the lookup in one message. chat-completion-agentic adds the loop: the model reads what a tool returned, decides whether to call another, and then answers.

    Your data goes in as OpenAI-style messages lists, and 10 to 50 seed conversations are usually enough.

    Read the docs →
  7. Upload training data as complete conversations

    You can now upload training data to distil labs as complete conversations, in the same OpenAI-style messages format the rest of your LLM stack already speaks. Traces, test sets and minimal datasets upload as they are, with no splitting into inputs and outputs, and a multi-turn interaction stays one record.

    The old question and answer columns were the shape a model consumes during training, not the shape anyone actually has. Your data should live in the shape your tooling produces.

    Read the docs →

Try what's new

Get the CLI with one command, then follow the docs to train your first small model.

$ curl -fsSL https://distillabs.ai/install.sh | sh