I believe this started with Sonnet 3.7. Sonnet 3.5 was and still is a much more reliable or 'accurate' model compared to 3.7 or Gemini 2.5 Pro. Interestingly, Sonnet 3.5 was only in 13th place on the
@lmarena_aileaderboard when 3.7 was released.
Which is interesting, because at that point, it was by far the best coding model.
@GergelyOroszreported that Sonnet 3.5 was so much ahead of any other model for coding that GitHub Copilot probably lost market share by not offering it sooner. Basically, Microsoft, the main
So we had a model which was beating anything else in real-life coding tasks, yet had a terrible place on
@lmarena_ai. With Sonnet 3.7, it changed; Anthropic finally started to optimize for the
@lmarena_airace! And with it came a model which became a "drunk professor". This
Interestingly,
@kamathematicmade a benchmark for "accuracy", where, not surprisingly, Sonnet 3.5 beats Sonnet 3.7 and Gemini 2.0 Flash absolutely beats Gemini 2.5 Pro
