Settings

Theme

Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

danluu.com

55 points by luu · 48 comments

Reader

5 threads
cb321

Something Dan does not observe in his article (perhaps Jamie does elsewhere? edit: or even Dan elsewhere) is that the same problem which makes the memory latency benchmark unrealistic (or at least misleading) often impacts hash table lookup benchmarks as mentioned at https://github.com/c-blake/bu/blob/main/doc/memlat.md and probably many other benchmarks. Essentially, CPU work prediction/speculative execution has become so good that much care is often required to measure latency rather than reciprocal throughput. This all started in the 1990s (or probably earlier with Cray), but I guess there's been an ongoing educational failure/oversimplification tendency.

Of course, throughput at one "level of work" is often latency at a higher level (like command-to-command execution, for example). So, "what matters" might be throughput or "latency". It all depends. :-) People often fiddle with such semantics to market methods, products, ideas, ... and marketing is often at cross purposes with understanding.

dom96

This is great. I've been building my own model benchmark lately and it has indeed been so easy to mess up the scoring. It's simply much harder to come up with an algorithm that combines all your individual scores into something that isn't broken in some special circumstances. That's why I think many just start capping the results.

jbellis

While I'm generally sympathetic to the idea that the public coding benchmarks are inadequate, to the point that I've written my own in the past and will likely do so again, the complaint here is that "these tasks don't match what I do in my day job" which is ~always going to be the case. The hope with benchmarks is that you can capture properties that generalize, from examining performance against small set of tasks, and I do think that this is at least directionally true for well-designed evals.

(There's https://withspecific.com/benchmarks/real-swe but since the tasks are private we still don't really know what they're representative of.)

  • Zigurd

    I don't mean to make you write a dissertation but to say that AI benchmarks can "capture properties that generalize from examining performance against a small set of tasks" is a bald assertion. It's a hypothesis without a theory behind it.

    I can profile some code and then tell you where the slow parts are. Unless there's some explanatory power to an AI benchmark, it's a bit like benchmarking pillows visually.

  • menaerus

    I started using gemini with caution given the "much worse" benchmarking points it has gotten and still does but in practice there's very little evidence I found in comparison to claude models. It performs really well on non trivial tasks.

stephantul

The (relevant) segue into tires elevates this post so much. I don’t understand why, but it does

  • IsTom

    I'm a little bit confused about these tire claims as

    > down to 0C / 32 F (he didn't test colder conditions)

    I get that people live in different places, but that's a huge caveat. How's that winter if you're not below 0°C? That sounds like "winter tires are worse than all-seasons tires in winter if you exclude winter".

    • mplanchard

      Yes, this stuck out to me, too. I use winter tires in the winter specifically for the snow and consistently below-freezing temps. Are there places people bother with winter tires that aren’t actually cold in the winter?

      • menaerus

        If warranted by the law then yes, you have to abide to it.

        • mplanchard

          Are there any places that require winter tires by law that don’t have cold winters? I don’t know of anywhere in the US that even requires winter tires, although I think some states require you to have winter tires OR chains.

          • RoddaWallPro

            There are mountain passes in Oregon with posted signs during winter that say "winter tires or chains required". I'm not sure if you can actually be cited by an officer for not having them, but certainly you are a hazard to everyone on the road if you don't have them :)

          • menaerus

            Sometimes the winters may be cold, some other times they may be exceptionally warm, and sometimes they're a mix of both. The law remains the same regardless what the winter was or is like. Countries that have exceptionally warm winters do not really have a winter conditions so I'd guess their law wouldn't have a requirement for winter tires or chains

            • mplanchard

              Yes of course, anyone who lives somewhere with winter knows that some winters are colder than others. In the late fall when it is time to swap the tires, you don’t know in advance whether the winter will be harsh or mild, but people put snow tires on anyway in places where harsh winters are common.

              It would only make sense to require them every winter if the large majority of winters were quite cold. As such, nowhere in the US requires winter tires, but someone else noted that Quebec does. This makes sense, since even a mild Quebec winter will spend most of the time getting below freezing.

              I would be very surprised if there were a place that required winter tires by law where the winter does not always get and stay at or below freezing for significant periods of time almost every winter.

          • auxym

            Quebec, Canada mandates winter tires between December 1st and March 15th.

            (But the winters are cold).

            • mplanchard

              Didn’t know it was required, but it makes sense, as it is indeed quite cold. Good thing I always have winter tires on in the winter anyway, so I’m not breaking the law when I drive to Montreal!

    • pm215

      I agree that testing colder conditions would be useful, but I guess you get into problems with it being ice and snow that you're testing, not the "summer tyres get too hard in the cold" hypothesis.

      It's winter if it's noticeably cooler than summer. For instance London's climate has a January average minimum temperature of just under 3C, which is very clearly winter compared to the July minimum of 14C. (figures from Met office site on their Heathrow station.) While the temperature does drop below 0C sometimes, it is not consistently below zero.

      In southern England pretty much nobody changes tyres for winter -- you just use the same set all year. Optimising for "2C in the wet" seems about right...

      • IsTom

        > but I guess you get into problems with it being ice and snow that you're testing

        If it's not actively snowing for a few days roads get clear by thawing during day because of salt/cars being hot/sun shining.

        The "summer tyres get too hard in the cold" idea could very well come from places where the typical winter is at least < -5°C. It's not uncommon in central/eastern/northern europe for temperatures to be < -10°C for extended periods of time.

        • rjsw

          The crossover point where winter tyres work better is < +7°C, it doesn't need to be below freezing.

          • pm215

            The linked article claims "in dry conditions, summer tires have the best grip down to 0C", though, which doesn't seem to match where you suggest the crossover point is.

            • attila-lendvai

              which is my experience, too.

              winter tires, especially the snow versions, are straight out a source of danger here in central europe.

              i just put some snow chains in for the winter surptises, and ride with my summer tires because its grip is clearly better on dry and wet bithumen, even in near zero temps, which is most of the season.

          • sgerenser

            Did you read the article? I thought his whole point is that little piece of folk wisdom didn't turn out to be true in actual testing.

    • 27183

      Lately I've just been leaving my snow tires on year round (Bridgestone Blizzaks). It gets cold here, between November and April the number of days with daytime high above freezing is small. I used to run all seasons in the summer and snows in the winter but they were lasting too long that way. It's better to wear them out within ~5yr.

      • threetonesun

        I've done that with "mild" Winter tires, really the thing full Summer tires excell at is removing water, which requires a center tread that's useless in snow. So if you don't get a lot of rain and it doesn't get too hot (some Winter compounds will get mushy and wear very quickly) they're fine. But really I've just described a Winter rated All-Season at this point, and you should buy those.

        • 27183

          All season tires are useless on snow and ice unless they're brand new. They're OK the first winter, after that it's bad news. On heavier vehicles I run BFGoodrich All Terrain T/A tires year round. They have good siped treads which grip on ice. For the car, studless Blizzaks have been holding up well year round. No abnormal wear so far, and they do fine in the wet, dry, heat, etc. I probably only drive 1-2 days/yr over 90°F though, in a hotter climate they might not be so great.

          [edit] obviously this is a tradeoff--I'm trading slightly reduced hot/dry/wet performance for massively increased winter performance. The reduction in summer performance is small enough to not be noticeable, whereas the increase in winter performance is large. On all season tires I would have to chain up a dozen or so times per year, often just to move the car like 3 car lengths out of a parking spot. I've only ever had to put the chains on once with snow tires on the car, and that was bashing through 6" of unplowed crusty icy stuff up a steep driveway.

          • dgacmu

            I dunno about this - I leave my crossclimate 2 aw's on year round and they're pretty fantastic in snow. We had a pretty good winter this year (44" from dec-feb) and they just worked. I didn't go meandering around any mountains on them, mind you, but they're so far ahead of most typical all-season tires it's almost not fair to compare. For -most- lazy people who don't want to swap tires or rims, they seem a better option than leaving true winter tires on year round. Obviously, there are exceptions depending on where you live.

            • 27183

              I've got two sets of wheels, and I rotate my tires every oil change, so it's not laziness that's the motivation. Instead, it was that with 2 sets of tires after 7 years of driving both sets still had tons of tread left but were starting to dry rot. I put about 60k miles on the car during those 7 years. So if I can get ~30k miles out of a single set of snows, and use them up within 5 years, that would be a much better use of resources. The cost isn't really a factor either, I just hate wasting stuff.

              I've currently got something like 15k mi on the tires and 2.5yr of year round use. They appear to be wearing evenly and normally.

    • TacticalCoder

      > How's that winter if you're not below 0°C?

      Seasons are called summer, autumn, winter and spring and each last three months and that bears no relation to whether or not there's snow or sub-zero temperatures.

      In the country I live in atm the rules regarding winter tires are not the absolute dumbest but they're still very dumb: you need to have either winter tires or all-seasons tires "if the conditions are winter'y". Which means, basically, both sub-zero AND either wet or icy. Sub-zero and all sunny means winter tires aren't mandatory. The reason it's still dumb it's that that correspond to, at most, 10 days per year. And this forces a lot of people to have worse performing tires during much more than 10 days. Which is probably the cause for a lot of accidents (e.g. people on days where it's + 3 C would be safer with summer tires, that do perform way better than "I've got winter tires because tomorrow at 7am it may or may not be -1 C and it may or may not be raining").

      Not that's of course dumbtardation but there's worse: there are countries where from that month to that month of winter, no matter the temperature, you must have winter (or all-seasons) tires. And at times you'll have an entire winter without freezing temperatures.

      So politicians who voted these laws are basically creating more accidents due to cars having inferior tires (the tires lobby does love it though).

      It's sad but it's how it is.

      • kuerbel

        A lot of drivers don't want it to be true but the quality of the tires absolutely matters. E.g. the usual ADAC test winner are actually good in dry conditions, but they are also expensive. The worst are just bad, but cheap. You get what you pay for. The best winter tires outclass bad summer tires, even up to lower middle end.

        The worst tires however are all-seasons.

  • vintagedave

    Maybe that it's an example of the same pattern in a real-world, non-digital area? It makes it feel more grounded.

    • daveguy

      It's something an LLM would never be able to come up with because it requires a depth of experience that LLMs just don't have. It's authentically human.

jdw64

I feel like AI development is hitting a wall now. Evaluations are driven by benchmarks, but I have no idea who is even evaluating the validity of these benchmarks. The idea that continuously training and scaling models will automatically yield broad capabilities across multiple domains hardly seems true anymore.

Benchmarks like Senior SWE-Bench apply discontinuous and arbitrary thresholds to evaluate results. If the generated code is semantically identical to the reference solution but even slightly exceeds a length cutoff, it fails? That seems like a genuinely flawed criterion, and the LLM-based graders themselves feel highly unstable.

Honestly, my impression is that LLM advancement is now completely dictated by benchmarks. But looking at recent trends, token prices are skyrocketing while coding capabilities haven't shown any massive improvements beyond a certain model generation. In fact, comparing GPT-6 Astra to 5.6 SOL, SOL often writes better code.

Considering all this, while training across diverse domains can pack a model with various pieces of knowledge, ultimately, it feels impossible for this approach to do something like derive the theory of relativity from medieval knowledge.

  • NitpickLawyer

    > AI development is hitting a wall now

    People have been saying this for at least 2 years now.

    > token prices are skyrocketing

    Today's SotA (fable and astra @ 50$ /Mtok output) are cheaper than o1-preview (sept '24, 60$ /Mtok output).

    And other models are workhorses, with much better capabilities, are at least 1 oom cheaper today than o1-preview. (I'm using this model, since it was the first "thinking" model)

    > it feels impossible for this approach to do something like

    The models have just provided lean proofs for FLT (a ~1M$ project that was expected to take a human expert ~5 years to complete) and a Millennium prize problem. These are current, relevant, and previously unsolved problems. The obsession people have with "proving relativity from stone-age data" is just moving the goalposts.

    • mjburgess

      No one spend ~10 mil USD using gpt 5.1 to solve a millenium prize problem with similar solutions in the training data. So we have no reference for whether 5.1, with that training data and that amount of moeny, could likewise produce such a solution.

      For each generation of advancement the "AI psychosis" of the previous wave wears off. Those who believed 4.6's reasoning was an accurate account of its behaviour; those who believed it had goals and solved useful problems reliably; and so on -- now, attribute only these things to Fable. And no doubt when Fable 6 comes along, it will be only v. 6 that does that.

      We have seen fairly marginal progress in LLM reliability and performance since the meaningful start of the high-inference/high-reasoning harness era. It just takes people a few model version bumps to break out of the addiction loop to realise this.

      At some point progress will stall entirely, and a couple years after that the spell will break and people will stop treating LLMs "as AI" in the wide-eyed sense, and start treating them as unreliable tools that map Text->Text -- as they do now with earlier model versions.

      • usef-

        There's definitely hype, but the newer models can undoubtedly do things the older ones couldn't. I had multiple long-term issues that earlier models couldn't solve (after repeated attempts) that fable did in 1 shot. On small models too, the differences in what I can trust them with has dramatically changed compared to a few months ago. Unless you breathlessly never touched the limitatioms of earlier models, it's very clear that the wall of limitations has been moving outward.

        Note also when a new benchmark is released, older models do worse on it than recent models, despite none of them being trained to the benchmark.

        • mjburgess

          Sure, because those businesses collect training data from users who are working on those problems. Perhaps this will never saturate and frontier users will always provide data to fill last generations gaps.

          My sense is the economics of that are going to collapse. It's currently extremely expensive to be on this endless retrain and inference cycle in order just to bake in additional marginal features.

          Maybe, maybe not. However I don't personally see anything other than 'one more leap', which might in any case arise from better integration with harnesses. I can foresee a step change due to harness reinforcement -- but other than that long mild refinements that are very expensive to acquire

          • usef-

            What tasks are you trying it with? It sounds like whatever you're doing isn't saturating the existing intelligence

  • lh712

    Nitpicking: It is not possible to derive the theory of relativity from medieval knowledge. At least not in any way resembling how the development actually happened, that is, heavily influenced by observations (or rather an iteration of speculation and observation). [Unless one counts both Galilean relativity and some parts of the theory of electromagnetism as being contained in medieval knowledge.]

    Perhaps it would be possible if the classically observed reality turned out to be possible (consistent under some reasonable conditions) only as a consequences of some deeper, sufficiently determined theory; but that would go far beyond our current level of knowledge, at least as far as I can say.

    • wongarsu

      The earliest you could reasonably derive special relativity is probably 1881 with the first version of the Michelson–Morley experiment, which showed that speed of light is the same no matter how you move through space. Which, on the scale of discoveries, is a pretty short timeline to the 1905 publication of special relativity.

      Which doesn't really stop you from training an LLM on 1904 knowledge and having it derive special relativity. LLMs trained with knowledge cutoffs that far back is a fun exercise, and one I've also dabbled in in the past. But you don't have anywhere near enough training data to reach SotA levels of intelligence, even if you had the money for those training runs. And lobotomizing all modern information out of LLMs doesn't seem viable either. So I don't see how you could ever turn that into a viable benchmark

      • cb321

        This is something of an aside on your first paragraph and lh712's. The theory of special relativity (and field theory in general) can be "derived" (as done in e.g., Landau, Lifshitz _Classical Theory Of Fields_) without empirical observation using concepts known to Galileo/Newton (if they count as "Medieval") - just with different assumptions/ideas about inertial frames. If you assume there is any phenomenon with a fixed observed speed in all "inertial frames" then you get special relativity with Einstein's gestalt-switch. After that it is, like so much in physics, a matter of thinking of an experiment to distinguish what matches capital-N Nature best.

        Thinking otherwise mistakes a fundamental error that the only way for something to happen is how it historically happened. Reconciling observations forced relativity historically, but it could have been sussed out without those. That there were very fast but finite speed phenomena (which could motivate relativity reducing to Newtonian models) was seen in 1676 by Rømer with the speed of light and eclipses of Io, a Galilean moon of Jupiter. It probably could have been done with Galileo's telescopes in 1610. (That is just one example of a phenomenon with a very fast yet finite speed to suggest other experiments to test that "hypothetical relativity".)

        TBH, this all relates to teaching "physics without calculus" and that sort of thing. Is the most clear presentation/derivation that which mirrors the history of our muddled yet ever demuddling ideas or that which starts from our most thoroughly demuddled ideas, best notations, etc.? The momentum is certainly the historical approach, yet there are notable exceptions.

        • lh712

          Well, exactly, and I think that that is in line with what I originally wrote, and which is that you need the general relativity principle (by general I don't mean the general relativity but the concept of the equivalence of inertial frames) and the theory electromagnetism. (By the way, I am familiar with the Landau--Lifshitz textbook.)

          (1) The relativity principle is exactly what you refer to yourself. You can't apply "Landau--Lifshitz"-like arguments without it. (And I don't think these principles count as medieval knowledge.)

          (2) I mentioned electromagnetism, because you need some clue for the concept of an absolute speed that is same for all inertial observers. This is very counterintuitive from our everyday experience, and counterintuitive from the point of view of somebody living in the time of "Galilean" or Newtonian mechanics. Theory of electromagnetism is the only thing that I am aware of, that is nearly (by a stretch) accessible at the level of "medieval" knowledge, from which the concept of constant speed follows. (Historically: Maxwell's completion of previously inconsistent equations of electromagnetism yielded a wave solution propagating with the constant speed of light. This was interpreted in terms of the ether originally, but it is a strong hint in itself for Einsteinian principle of relativity; and a reasoning along the lines you alluded to can them be applied, at least in principle.)

          • cb321

            Yeah. My point was mostly to expand upon your original post - sorry if it sounded like I disagreed. You say the same in your original "iteration of speculation and observation". Everything else is about what "derive" and "knowledge" might mean (I agree "medieval" usually means pre-Renaissance and Galileo is modern-era), how much empirical "proof" is in "derive", etc. However, all you need for "possible" is "the idea", some "consequences", and ways to test.

            So, to push back a tad on your more recent "only thing that I am aware of" and to maybe explain my above point better, I do think there was enough information/ideas in the abstract in Galileo/Newton/Leibniz' times to suggest the idea. Leibniz himself pushed back hard on Newton's absolute space/time (long before Mach). For Leibniz, it would have been counterintuitive space/time vs. counterintuitive fixed speed. So, that pushback itself could have been enough of a "clue" in your terms -- in some alternate timeline -- to drive a speculation-observation cycle starting from different inertial concepts - with Galilean moon eclipses then enough a clue that fast speeds existed to fool our slow-speed intuitions. (And all this in the "modern era".)

            So, I continue to think it "not impossible" that the relativity principle could have arisen before any EM theory at all -- it just didn't. That matters for these kinds of speculative questions about what information horizons support what developments. A single "fast enough" fixed speed is is not that wild & crazy. The modern world has a zillion obscure physics theories like that (mostly just because so many more people work on that stuff, but that's a probability thing, not a possibility thing, and I suppose also partly inspired by how physics turned out - so not truly independent). History is littered with things that could have happened, but didn't.

            P.S.: and apologies for "mistakes a fundamental error". I of course meant "makes a", if you wanted any evidence that I was not an LLM. ;-)

            • lh712

              Thank you for your reply! Yes, I fully agree that meaning of "derive" and other related concepts are non-trivial. (And I actually changed my own "working definition" between my two posts; while in the first one I was leaning more towards theories anchored by observation, even if indirectly, in the second one I slipped into more mathematical/philosophical perspective. The second part of my first reply makes sense only from the first kind of perspective, because in that part I imagine that the observed reality, even at the medieval level of knowledge, might perhaps necessitate a particular kind of underlying theory, from which the relativity principle would then "fall out". That was the only way I could think of at that moment how SR could have been "derived" without further observations.)

              Your points about Leibniz are very interesting, and while I find that line of thought fascinating I must admit my own lack of knowledge of the historical context. In any case, thank you for pointing out that idea!

              • cb321

                You're welcome.

                Another lesser known wrinkle along these lines is that if more of Aristarchus of Samos' work on heliocentricity had been developed into a "calculational framework" for planetary motion by a contemporaneous Ptolemy competitor (a la Copernicus in 1543), various relativity ideas might well have started in 270BC instead of 1600 AD.

                Aristarchus was already WAY ahead of a few games - making a guess that "absolute rest" was illusory and tiny stars were similar to giant Sol. If not for Archimedes' serious star power, we might not even know of Aristarchus' heliocentric ideas! Relativity ideas of various kinds are short intellectual hops from "the absolute rest of your intuition is illusory".

                Aristarchus and his supporters just had to assert The Stars were very distant -- far enough for there to be NO PARALLAX! Various "coincidences" all conspired to suppress Aristarchus' plausibility, like: A) Just HOW DISTANT stars are/local galactic stellar density &| B) poor human VISUAL ACUITY relative to C) Earth ORBIT vs. Sun LUMINOSITY (parallax baseline) & maybe existence of Moon to vent atmo &| D) VERY SLOW (3 millennium) development from near prehistoric glass (1500BC) to grinding lenses for human eyes in the 1300s AD leading to Galileo's telescopes & etc. If ANY of (A)..(D) were about ~10x better (ALL of which are imaginably so), the day that universe changed could've been much earlier. Heck, even if Aristarchus just had a really good spin doctor like a major religion pushing the plausibility of "Sun=A close Star", that might've been enough.

                Archimedes almost invented the underlying ideas of integration and limits and all that, too. Close but no cigar, but (had that been in hand) analytic geometry and differential equations are short steps away. I mean, Newton surely noted the equivalence of gravitational and inertial "mass" (charge vs. kinematics) which is the weak principle of equivalence of GR, after all. Like I said - "littered". ;-) There are probably whole books written on the topic (or adjacent topics) of "All the Things Humanity Nearly Figured Out Earlier" that have a more historical bent than the usual "sci-fi tilt".

  • WarmWash

    Math produces a bunch of theories when you extrapolate a system forward (or sometimes backwards)

    Experimentation is the hammer that smashes all the incorrect theories.

    Without the ability to do experimentation, deriving new laws is virtually impossible. A new next step for AI would be coming up with novel experiments, because that is often the hardest part.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection