Live data from Hacker News

Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

danluu.com

11–20 of 47 posts

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#11
Something Dan does not observe in his article (perhaps Jamie does elsewhere? edit: or even Dan elsewhere) is that the same problem which makes the memory latency benchmark unrealistic (or at least misleading) often impacts hash table lookup benchmarks as mentioned at https://github.com/c-blake/bu/blob/main/doc/memlat.md and probably many other benchmarks. Essentially, CPU work prediction/speculative execution has become so good that much care is often required to measure latency rather than reciprocal throughput. This all started in the 1990s (or probably earlier with Cray), but I guess there's been an ongoing educational failure/oversimplification tendency.

Of course, throughput at one "level of work" is often latency at a higher level (like command-to-command execution, for example). So, "what matters" might be throughput or "latency". It all depends. :-) People often fiddle with such semantics to market methods, products, ideas, ... and marketing is often at cross purposes with understanding.

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#12
post #9
post #7

I feel like AI development is hitting a wall now. Evaluations are driven by benchmarks, but I have no idea who is even evaluating the validity of these benchmarks. The idea that continuously training and scaling models will automatically yield broad capabilities across multiple domains hardly seems true anymore. Benchmarks like Senior SWE-Bench apply discontinuous and arbitrary thresholds to evaluate results. If the…

Nitpicking: It is not possible to derive the theory of relativity from medieval knowledge. At least not in any way resembling how the development actually happened, that is, heavily influenced by observations (or rather an iteration of speculation and observation). [Unless one counts both Galilean relativity and some parts of the theory of electromagnetism as being contained in medieval knowledge.] Perhaps it would b…

The earliest you could reasonably derive special relativity is probably 1881 with the first version of the Michelson–Morley experiment, which showed that speed of light is the same no matter how you move through space. Which, on the scale of discoveries, is a pretty short timeline to the 1905 publication of special relativity.

Which doesn't really stop you from training an LLM on 1904 knowledge and having it derive special relativity. LLMs trained with knowledge cutoffs that far back is a fun exercise, and one I've also dabbled in in the past. But you don't have anywhere near enough training data to reach SotA levels of intelligence, even if you had the money for those training runs. And lobotomizing all modern information out of LLMs doesn't seem viable either. So I don't see how you could ever turn that into a viable benchmark

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#13
post #7

I feel like AI development is hitting a wall now. Evaluations are driven by benchmarks, but I have no idea who is even evaluating the validity of these benchmarks. The idea that continuously training and scaling models will automatically yield broad capabilities across multiple domains hardly seems true anymore. Benchmarks like Senior SWE-Bench apply discontinuous and arbitrary thresholds to evaluate results. If the…

> AI development is hitting a wall now

People have been saying this for at least 2 years now.

> token prices are skyrocketing

Today's SotA (fable and astra @ 50$ /Mtok output) are cheaper than o1-preview (sept '24, 60$ /Mtok output).

And other models are workhorses, with much better capabilities, are at least 1 oom cheaper today than o1-preview. (I'm using this model, since it was the first "thinking" model)

> it feels impossible for this approach to do something like

The models have just provided lean proofs for FLT (a ~1M$ project that was expected to take a human expert ~5 years to complete) and a Millennium prize problem. These are current, relevant, and previously unsolved problems. The obsession people have with "proving relativity from stone-age data" is just moving the goalposts.

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#14
post #9

Earlier quoted context omitted.

Nitpicking: It is not possible to derive the theory of relativity from medieval knowledge. At least not in any way resembling how the development actually happened, that is, heavily influenced by observations (or rather an iteration of speculation and observation). [Unless one counts both Galilean relativity and some parts of the theory of electromagnetism as being contained in medieval knowledge.] Perhaps it would b…

The earliest you could reasonably derive special relativity is probably 1881 with the first version of the Michelson–Morley experiment, which showed that speed of light is the same no matter how you move through space. Which, on the scale of discoveries, is a pretty short timeline to the 1905 publication of special relativity. Which doesn't really stop you from training an LLM on 1904 knowledge and having it derive s…

This is something of an aside on your first paragraph and lh712's. The theory of special relativity (and field theory in general) can be "derived" (as done in e.g., Landau, Lifshitz _Classical Theory Of Fields_) without empirical observation using concepts known to Galileo/Newton (if they count as "Medieval") - just with different assumptions/ideas about inertial frames. If you assume there is any phenomenon with a fixed observed speed in all "inertial frames" then you get special relativity with Einstein's gestalt-switch. After that it is, like so much in physics, a matter of thinking of an experiment to distinguish what matches capital-N Nature best.

Thinking otherwise mistakes a fundamental error that the only way for something to happen is how it historically happened. Reconciling observations forced relativity historically, but it could have been sussed out without those. That there were very fast but finite speed phenomena (which could motivate relativity reducing to Newtonian models) was seen in 1676 by Rømer with the speed of light and eclipses of Io, a Galilean moon of Jupiter. It probably could have been done with Galileo's telescopes in 1610. (That is just one example of a phenomenon with a very fast yet finite speed to suggest other experiments to test that "hypothetical relativity".)

TBH, this all relates to teaching "physics without calculus" and that sort of thing. Is the most clear presentation/derivation that which mirrors the history of our muddled yet ever demuddling ideas or that which starts from our most thoroughly demuddled ideas, best notations, etc.? The momentum is certainly the historical approach, yet there are notable exceptions.

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#15
post #7

I feel like AI development is hitting a wall now. Evaluations are driven by benchmarks, but I have no idea who is even evaluating the validity of these benchmarks. The idea that continuously training and scaling models will automatically yield broad capabilities across multiple domains hardly seems true anymore. Benchmarks like Senior SWE-Bench apply discontinuous and arbitrary thresholds to evaluate results. If the…

> AI development is hitting a wall now People have been saying this for at least 2 years now. > token prices are skyrocketing Today's SotA (fable and astra @ 50$ /Mtok output) are cheaper than o1-preview (sept '24, 60$ /Mtok output). And other models are workhorses, with much better capabilities, are at least 1 oom cheaper today than o1-preview. (I'm using this model, since it was the first "thinking" model) > it fee…

No one spend ~10 mil USD using gpt 5.1 to solve a millenium prize problem with similar solutions in the training data. So we have no reference for whether 5.1, with that training data and that amount of moeny, could likewise produce such a solution.

For each generation of advancement the "AI psychosis" of the previous wave wears off. Those who believed 4.6's reasoning was an accurate account of its behaviour; those who believed it had goals and solved useful problems reliably; and so on -- now, attribute only these things to Fable. And no doubt when Fable 6 comes along, it will be only v. 6 that does that.

We have seen fairly marginal progress in LLM reliability and performance since the meaningful start of the high-inference/high-reasoning harness era. It just takes people a few model version bumps to break out of the addiction loop to realise this.

At some point progress will stall entirely, and a couple years after that the spell will break and people will stop treating LLMs "as AI" in the wide-eyed sense, and start treating them as unreliable tools that map Text->Text -- as they do now with earlier model versions.

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#16
post #14

Earlier quoted context omitted.

The earliest you could reasonably derive special relativity is probably 1881 with the first version of the Michelson–Morley experiment, which showed that speed of light is the same no matter how you move through space. Which, on the scale of discoveries, is a pretty short timeline to the 1905 publication of special relativity. Which doesn't really stop you from training an LLM on 1904 knowledge and having it derive s…

This is something of an aside on your first paragraph and lh712's. The theory of special relativity (and field theory in general) can be "derived" (as done in e.g., Landau, Lifshitz _Classical Theory Of Fields_) without empirical observation using concepts known to Galileo/Newton (if they count as "Medieval") - just with different assumptions/ideas about inertial frames. If you assume there is any phenomenon with a f…

Well, exactly, and I think that that is in line with what I originally wrote, and which is that you need the general relativity principle (by general I don't mean the general relativity but the concept of the equivalence of inertial frames) and the theory electromagnetism. (By the way, I am familiar with the Landau--Lifshitz textbook.)

(1) The relativity principle is exactly what you refer to yourself. You can't apply "Landau--Lifshitz"-like arguments without it. (And I don't think these principles count as medieval knowledge.)

(2) I mentioned electromagnetism, because you need some clue for the concept of an absolute speed that is same for all inertial observers. This is very counterintuitive from our everyday experience, and counterintuitive from the point of view of somebody living in the time of "Galilean" or Newtonian mechanics. Theory of electromagnetism is the only thing that I am aware of, that is nearly (by a stretch) accessible at the level of "medieval" knowledge, from which the concept of constant speed follows. (Historically: Maxwell's completion of previously inconsistent equations of electromagnetism yielded a wave solution propagating with the constant speed of light. This was interpreted in terms of the ether originally, but it is a strong hint in itself for Einsteinian principle of relativity; and a reasoning along the lines you alluded to can them be applied, at least in principle.)

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#17
post #10
post #6

Earlier quoted context omitted.

I'm a little bit confused about these tire claims as > down to 0C / 32 F (he didn't test colder conditions) I get that people live in different places, but that's a huge caveat. How's that winter if you're not below 0°C? That sounds like "winter tires are worse than all-seasons tires in winter if you exclude winter".

I agree that testing colder conditions would be useful, but I guess you get into problems with it being ice and snow that you're testing, not the "summer tyres get too hard in the cold" hypothesis. It's winter if it's noticeably cooler than summer. For instance London's climate has a January average minimum temperature of just under 3C, which is very clearly winter compared to the July minimum of 14C. (figures from M…

> but I guess you get into problems with it being ice and snow that you're testing

If it's not actively snowing for a few days roads get clear by thawing during day because of salt/cars being hot/sun shining.

The "summer tyres get too hard in the cold" idea could very well come from places where the typical winter is at least < -5°C. It's not uncommon in central/eastern/northern europe for temperatures to be < -10°C for extended periods of time.

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#18
While I'm generally sympathetic to the idea that the public coding benchmarks are inadequate, to the point that I've written my own in the past and will likely do so again, the complaint here is that "these tasks don't match what I do in my day job" which is ~always going to be the case. The hope with benchmarks is that you can capture properties that generalize, from examining performance against small set of tasks, and I do think that this is at least directionally true for well-designed evals.

(There's https://withspecific.com/benchmarks/real-swe but since the tasks are private we still don't really know what they're representative of.)

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#19
post #7

I feel like AI development is hitting a wall now. Evaluations are driven by benchmarks, but I have no idea who is even evaluating the validity of these benchmarks. The idea that continuously training and scaling models will automatically yield broad capabilities across multiple domains hardly seems true anymore. Benchmarks like Senior SWE-Bench apply discontinuous and arbitrary thresholds to evaluate results. If the…

[dead]

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#20

Earlier quoted context omitted.

> AI development is hitting a wall now People have been saying this for at least 2 years now. > token prices are skyrocketing Today's SotA (fable and astra @ 50$ /Mtok output) are cheaper than o1-preview (sept '24, 60$ /Mtok output). And other models are workhorses, with much better capabilities, are at least 1 oom cheaper today than o1-preview. (I'm using this model, since it was the first "thinking" model) > it fee…

No one spend ~10 mil USD using gpt 5.1 to solve a millenium prize problem with similar solutions in the training data. So we have no reference for whether 5.1, with that training data and that amount of moeny, could likewise produce such a solution. For each generation of advancement the "AI psychosis" of the previous wave wears off. Those who believed 4.6's reasoning was an accurate account of its behaviour; those w…

There's definitely hype, but the newer models can undoubtedly do things the older ones couldn't. I had multiple long-term issues that earlier models couldn't solve (after repeated attempts) that fable did in 1 shot. On small models too, the differences in what I can trust them with has dramatically changed compared to a few months ago. Unless you breathlessly never touched the limitatioms of earlier models, it's very clear that the wall of limitations has been moving outward.

Note also when a new benchmark is released, older models do worse on it than recent models, despite none of them being trained to the benchmark.

Post reply on HN