Consumer apps
For consumer apps where every user interaction calls a model.
Cost scales with your growing user count
Every new user adds requests, competition caps what you can charge, and gross margin pays the difference. The model bill grows into one of your largest variable costs.
At hundreds of millions of requests a month, a few cents per thousand requests decide whether a feature is viable at all.
Cheaper tiers make your users churn
Dropping to your provider's smaller tier is the one lever that moves the bill, and it's the one that costs you accuracy. Wrong answers come back as support tickets, abandoned sessions and cancelled subscriptions.
The safe alternatives barely register. Trimming prompts and adding caching are afternoon projects worth single-digit percentages, and routers shuffle the same tokens onto a different bill.
Custom SLM reclaims your margins without churn
We train a specialized SLM on your production workload and serve it behind the OpenAI-compatible endpoint you already call. On bounded, high-volume tasks it clears your quality bar from a model roughly 100x smaller.
Knowunity's accuracy went up, from 81% to 93%, while the bill fell 68%. The model is yours, so cost per successful task tracks your traffic rather than your provider's price list.
Where it applies
- Request classification and routing. Label every incoming request so it reaches the right workflow. This is the Knowunity workload below, running at over 100M requests a month.
- Content moderation. Flag user-generated text against your policy, on every post rather than a sample.
- Structured extraction. Turn free-text user input into the fields your product actually stores.
- Search and query understanding. Rewrite and classify queries before they hit your index.
- Bounded-domain chat. Tutoring, support and guided flows where the subject matter is fixed and the answer format is known.
- Personalized notifications and summaries. High-volume generation per user, where a frontier model per send is what breaks the economics.
In production
Knowunity
Knowunity classifies every incoming student request to route it to the right workflow, running on distil labs models at hundreds of millions of requests per month.
68%
lower inference costs
81% → 93%
task accuracy, versus the LLM it replaced
100M+
requests per month served
Read the Knowunity case study →“Using distil labs, we were able to spin up highly accurate custom small models tailored to our workflows in no time. Those models cut our inference costs by 68% without sacrificing quality.”
Find out what your workload should cost
We identify where you are overspending, evaluate the optimizations available, and show you the lowest-cost configuration that meets your quality bar. It starts with an export or one day of traffic.