Live data from Hacker News

Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

danluu.com

21–30 of 46 posts

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#21
post #20

Earlier quoted context omitted.

No one spend ~10 mil USD using gpt 5.1 to solve a millenium prize problem with similar solutions in the training data. So we have no reference for whether 5.1, with that training data and that amount of moeny, could likewise produce such a solution. For each generation of advancement the "AI psychosis" of the previous wave wears off. Those who believed 4.6's reasoning was an accurate account of its behaviour; those w…

There's definitely hype, but the newer models can undoubtedly do things the older ones couldn't. I had multiple long-term issues that earlier models couldn't solve (after repeated attempts) that fable did in 1 shot. On small models too, the differences in what I can trust them with has dramatically changed compared to a few months ago. Unless you breathlessly never touched the limitatioms of earlier models, it's very…

Sure, because those businesses collect training data from users who are working on those problems. Perhaps this will never saturate and frontier users will always provide data to fill last generations gaps.

My sense is the economics of that are going to collapse. It's currently extremely expensive to be on this endless retrain and inference cycle in order just to bake in additional marginal features.

Maybe, maybe not. However I don't personally see anything other than 'one more leap', which might in any case arise from better integration with harnesses. I can foresee a step change due to harness reinforcement -- but other than that long mild refinements that are very expensive to acquire

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#22
post #6

The (relevant) segue into tires elevates this post so much. I don’t understand why, but it does

I'm a little bit confused about these tire claims as > down to 0C / 32 F (he didn't test colder conditions) I get that people live in different places, but that's a huge caveat. How's that winter if you're not below 0°C? That sounds like "winter tires are worse than all-seasons tires in winter if you exclude winter".

> How's that winter if you're not below 0°C?

Seasons are called summer, autumn, winter and spring and each last three months and that bears no relation to whether or not there's snow or sub-zero temperatures.

In the country I live in atm the rules regarding winter tires are not the absolute dumbest but they're still very dumb: you need to have either winter tires or all-seasons tires "if the conditions are winter'y". Which means, basically, both sub-zero AND either wet or icy. Sub-zero and all sunny means winter tires aren't mandatory. The reason it's still dumb it's that that correspond to, at most, 10 days per year. And this forces a lot of people to have worse performing tires during much more than 10 days. Which is probably the cause for a lot of accidents (e.g. people on days where it's + 3 C would be safer with summer tires, that do perform way better than "I've got winter tires because tomorrow at 7am it may or may not be -1 C and it may or may not be raining").

Not that's of course dumbtardation but there's worse: there are countries where from that month to that month of winter, no matter the temperature, you must have winter (or all-seasons) tires. And at times you'll have an entire winter without freezing temperatures.

So politicians who voted these laws are basically creating more accidents due to cars having inferior tires (the tires lobby does love it though).

It's sad but it's how it is.

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#23
post #6

Earlier quoted context omitted.

I'm a little bit confused about these tire claims as > down to 0C / 32 F (he didn't test colder conditions) I get that people live in different places, but that's a huge caveat. How's that winter if you're not below 0°C? That sounds like "winter tires are worse than all-seasons tires in winter if you exclude winter".

> How's that winter if you're not below 0°C? Seasons are called summer, autumn, winter and spring and each last three months and that bears no relation to whether or not there's snow or sub-zero temperatures. In the country I live in atm the rules regarding winter tires are not the absolute dumbest but they're still very dumb: you need to have either winter tires or all-seasons tires "if the conditions are winter'y"…

A lot of drivers don't want it to be true but the quality of the tires absolutely matters. E.g. the usual ADAC test winner are actually good in dry conditions, but they are also expensive. The worst are just bad, but cheap. You get what you pay for. The best winter tires outclass bad summer tires, even up to lower middle end.

The worst tires however are all-seasons.

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#24
post #16
post #14

Earlier quoted context omitted.

This is something of an aside on your first paragraph and lh712's. The theory of special relativity (and field theory in general) can be "derived" (as done in e.g., Landau, Lifshitz _Classical Theory Of Fields_) without empirical observation using concepts known to Galileo/Newton (if they count as "Medieval") - just with different assumptions/ideas about inertial frames. If you assume there is any phenomenon with a f…

Well, exactly, and I think that that is in line with what I originally wrote, and which is that you need the general relativity principle (by general I don't mean the general relativity but the concept of the equivalence of inertial frames) and the theory electromagnetism. (By the way, I am familiar with the Landau--Lifshitz textbook.) (1) The relativity principle is exactly what you refer to yourself. You can't appl…

Yeah. My point was mostly to expand upon your original post - sorry if it sounded like I disagreed. You say the same in your original "iteration of speculation and observation". Everything else is about what "derive" and "knowledge" might mean (I agree "medieval" usually means pre-Renaissance and Galileo is modern-era), how much empirical "proof" is in "derive", etc. However, all you need for "possible" is "the idea", some "consequences", and ways to test.

So, to push back a tad on your more recent "only thing that I am aware of" and to maybe explain my above point better, I do think there was enough information/ideas in the abstract in Galileo/Newton/Leibniz' times to suggest the idea. Leibniz himself pushed back hard on Newton's absolute space/time (long before Mach). For Leibniz, it would have been counterintuitive space/time vs. counterintuitive fixed speed. So, that pushback itself could have been enough of a "clue" in your terms -- in some alternate timeline -- to drive a speculation-observation cycle starting from different inertial concepts - with Galilean moon eclipses then enough a clue that fast speeds existed to fool our slow-speed intuitions. (And all this in the "modern era".)

So, I continue to think it "not impossible" that the relativity principle could have arisen before any EM theory at all -- it just didn't. That matters for these kinds of speculative questions about what information horizons support what developments. A single "fast enough" fixed speed is is not that wild & crazy. The modern world has a zillion obscure physics theories like that (mostly just because so many more people work on that stuff, but that's a probability thing, not a possibility thing, and I suppose also partly inspired by how physics turned out - so not truly independent). History is littered with things that could have happened, but didn't.

P.S.: and apologies for "mistakes a fundamental error". I of course meant "makes a", if you wanted any evidence that I was not an LLM. ;-)

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#25
post #6

The (relevant) segue into tires elevates this post so much. I don’t understand why, but it does

I'm a little bit confused about these tire claims as > down to 0C / 32 F (he didn't test colder conditions) I get that people live in different places, but that's a huge caveat. How's that winter if you're not below 0°C? That sounds like "winter tires are worse than all-seasons tires in winter if you exclude winter".

Lately I've just been leaving my snow tires on year round (Bridgestone Blizzaks). It gets cold here, between November and April the number of days with daytime high above freezing is small. I used to run all seasons in the summer and snows in the winter but they were lasting too long that way. It's better to wear them out within ~5yr.

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#26

The (relevant) segue into tires elevates this post so much. I don’t understand why, but it does

Maybe that it's an example of the same pattern in a real-world, non-digital area? It makes it feel more grounded.

It's something an LLM would never be able to come up with because it requires a depth of experience that LLMs just don't have. It's authentically human.

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#27
post #20

Earlier quoted context omitted.

There's definitely hype, but the newer models can undoubtedly do things the older ones couldn't. I had multiple long-term issues that earlier models couldn't solve (after repeated attempts) that fable did in 1 shot. On small models too, the differences in what I can trust them with has dramatically changed compared to a few months ago. Unless you breathlessly never touched the limitatioms of earlier models, it's very…

Sure, because those businesses collect training data from users who are working on those problems. Perhaps this will never saturate and frontier users will always provide data to fill last generations gaps. My sense is the economics of that are going to collapse. It's currently extremely expensive to be on this endless retrain and inference cycle in order just to bake in additional marginal features. Maybe, maybe not…

What tasks are you trying it with? It sounds like whatever you're doing isn't saturating the existing intelligence

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#28
post #25
post #6

Earlier quoted context omitted.

I'm a little bit confused about these tire claims as > down to 0C / 32 F (he didn't test colder conditions) I get that people live in different places, but that's a huge caveat. How's that winter if you're not below 0°C? That sounds like "winter tires are worse than all-seasons tires in winter if you exclude winter".

Lately I've just been leaving my snow tires on year round (Bridgestone Blizzaks). It gets cold here, between November and April the number of days with daytime high above freezing is small. I used to run all seasons in the summer and snows in the winter but they were lasting too long that way. It's better to wear them out within ~5yr.

I've done that with "mild" Winter tires, really the thing full Summer tires excell at is removing water, which requires a center tread that's useless in snow. So if you don't get a lot of rain and it doesn't get too hot (some Winter compounds will get mushy and wear very quickly) they're fine. But really I've just described a Winter rated All-Season at this point, and you should buy those.

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#29
post #17
post #10

Earlier quoted context omitted.

I agree that testing colder conditions would be useful, but I guess you get into problems with it being ice and snow that you're testing, not the "summer tyres get too hard in the cold" hypothesis. It's winter if it's noticeably cooler than summer. For instance London's climate has a January average minimum temperature of just under 3C, which is very clearly winter compared to the July minimum of 14C. (figures from M…

> but I guess you get into problems with it being ice and snow that you're testing If it's not actively snowing for a few days roads get clear by thawing during day because of salt/cars being hot/sun shining. The "summer tyres get too hard in the cold" idea could very well come from places where the typical winter is at least < -5°C. It's not uncommon in central/eastern/northern europe for temperatures to be < -10°C…

The crossover point where winter tyres work better is < +7°C, it doesn't need to be below freezing.

Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

#30
post #18

While I'm generally sympathetic to the idea that the public coding benchmarks are inadequate, to the point that I've written my own in the past and will likely do so again, the complaint here is that "these tasks don't match what I do in my day job" which is ~always going to be the case. The hope with benchmarks is that you can capture properties that generalize, from examining performance against small set of tasks,…

I don't mean to make you write a dissertation but to say that AI benchmarks can "capture properties that generalize from examining performance against a small set of tasks" is a bald assertion. It's a hypothesis without a theory behind it.

I can profile some code and then tell you where the slow parts are. Unless there's some explanatory power to an AI benchmark, it's a bit like benchmarking pillows visually.

Post reply on HN