Model Fatigue is Real

· Moe Khalil ·

4 min read Original article ↗

Two days ago, Anthropic released Fable 5.1 and Mythos 5.1. Their launch post says that “they’re the world’s most advanced models for coding and knowledge work”. Today, OpenAI released GPT-6 Astra. Their launch post claims that it’s “the world’s most intelligent and aligned model”. In the meantime, Gemini released 3.8 Flash, which they claim will “deliver next-generation intelligence”, and Meta released Muse Spark 1.3, their “most powerful frontier AI model yet”.

If you’re not overwhelmed yet, there’s more. XAI, Deepseek, Z.ai, Qwen, NVIDIA, and Tencent have all released models in the past month. Each one claims they’re either “the best”, “the best at X”, or “the most efficient”. They make those claims based on benchmarks that they somehow all claim to top.

Even as someone who spends the majority of their time wrangling LLMs, it’s exhausting. I’m already busy enough trying to get my job done, and now on top of that I have to become an expert on what’s coming out lest I fail to keep up with people who do. Obviously, I want to be as productive as possible. But at this point chasing the “perfect model” is wasting more of my time than it’s giving me back. What’s more, I know for a fact that I don’t need Fable 5.1 for the majority of what I’m doing on a day-to-day basis. But I’m lazy, and I don’t want to fall behind, so I just set the default to whatever feels smartest and “let it cook”.

The issue is, for the model companies, this laziness is a feature - not a bug. They prefer that I use Sol 5.6 Ultra to summarize my email, and they have no incentive to tell me when a cheaper model would work just as well. Their job is to keep pushing the frontier. Figuring out when not to use the frontier is left to me.

Now, I’m not claiming that the frontier models aren’t incredible. I’ve dedicated my career to working on LLMs because I think they’re genuinely the greatest technology of my lifetime.

But it does get to a point.

At this point, I’d like to add a quick caveat. I recently joined LiteLLM as a product engineer where I’ve been working on their auto router. I was excited to join because this is a very real problem that I myself have felt, and the ideal version is not limited to a single family of models — which is one of LiteLLM’s strength.

It’s also a very interesting technical problem. Put simply, there is a minimum viable model for any given closed-ended task. Given a closed-ended task, if I were to run every model on that task, a subset would fail and another subset would succeed. Of those that succeed, there will be one that succeeds at the cheapest price. This extends to other parameters like latency, but I’ll focus on price for now.

So, we know that this “perfect model for the task exists”. The problem becomes, is there a way to figure out what that model is before the task has been run? That’s what we’re trying to figure out.

The problem becomes even more compelling when you try comparing models at different costs on real world tasks. Our initial intuition told us that if Terra and Sol could both solve a task, then Terra (which costs half as much according to the model card) would solve it for cheaper. But when we ran the tests, Sol was solving the more difficult tasks at a fraction of the price of the cheaper model. Because it was smarter, it would take a smarter approach, and would get to the finished product way faster.

Fascinating, right?

So how do we solve this? We’ve been testing multiple approaches. Our most recent approach (and most successful so far) involved using public benchmark data to infer what the cheapest model to solve a task would be from heuristics in the prompt. We’ve also experimented with LLM classification and hybrid approaches. All still imperfect, but a step up from nothing. The end goal is that I shouldn’t have to think about any of this. I should be able to get my work done, and trust that the “perfect model” is being used behind the scenes.

LiteLLM is an open-source AI gateway that puts your full AI stack behind one OpenAI-compatible key. Our Auto Router is still in beta, and we are working hard on improving it every day. Any and all feedback would be extremely helpful. You can reach me at moe@berri.ai

Discussion about this post

Ready for more?