leibler is a lab that helps you own your intelligence.
We post-train smaller models on your tasks until they clear your bar, then run them for you. Higher quality, lower bill.
# one line to integrate
import leibler
from openai import OpenAI
client = leibler.wrap(OpenAI())
00
The bar never moves. The model climbs to it.
Your frontier model sets the bar. A smaller model trains on its passing attempts until it clears reliably.
task ticket_classificationiteration 01clearing the bar 0%graduated 0 of 4
01Evaluate
Add the wrapper or send a week of traces. We rebuild your environment from them, verified against the traces, plus synthetic data. In 48 hours you know which tasks a smaller model already handles.
02Build
For the rest, we post-train a model on your traces until it clears the bar.
03Serve
We deploy it in production, tuned for your traffic. Our cloud, yours, a VPC, or on-prem, and on your own hardware with the provider you are comfortable with.
04Route
You apply the routing plan in your own config.
One customer's plan, graded against their own frontier model. In 48 hours you get the same for your tasks.
| task | clears the bar | cost vs frontier |
|---|---|---|
| summarize_ | 96% | 18× cheaper |
| ticket_ | 100% | 22× cheaper |
| extract_ | 98% | 15× cheaper |
| draft_ | 91% | 10× cheaper |
Nothing ships below the bar. Every release is re-graded against it first.
03
Your data stays yours.
Your traces train only your models, are never pooled, and are deleted on request. Capture can't slow or break your calls.
Kullback, the harness that rebuilds your environment from traces and checks the rebuild by replay, is Apache-2.0 on GitHub. What it builds, what is measured, and the one claim it is out to earn.
What happens to my data?
Your traces train only your models. Nothing is pooled across customers, and everything is deleted on request.
Who owns the fine-tuned weights?
You do. The models are built from your traces, for you, and deploy on your infrastructure or ours.
What if quality drops below the bar?
It doesn't ship. Every release is graded against the bar first; anything below it is held back.
How is this different from just using a cheaper model?
On its own it doesn't clear your bar. We prove on your traces which tasks it can take before anything ships.
What does it cost?
Bring one task. In thirty minutes we'll tell you what it takes and whether a smaller model clears the bar.
Bring one task.
Thirty minutes. We tell you whether a smaller model can clear the bar on it, and what it takes.
Krrish Agarwalla · post-training research at Stanford and LMU Munich