Of course, throughput at one "level of work" is often latency at a higher level (like command-to-command execution, for example). So, "what matters" might be throughput or "latency". It all depends. :-) People often fiddle with such semantics to market methods, products, ideas, ... and marketing is often at cross purposes with understanding.
Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
11–20 of 45 posts
Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#12I feel like AI development is hitting a wall now. Evaluations are driven by benchmarks, but I have no idea who is even evaluating the validity of these benchmarks. The idea that continuously training and scaling models will automatically yield broad capabilities across multiple domains hardly seems true anymore. Benchmarks like Senior SWE-Bench apply discontinuous and arbitrary thresholds to evaluate results. If the…
Nitpicking: It is not possible to derive the theory of relativity from medieval knowledge. At least not in any way resembling how the development actually happened, that is, heavily influenced by observations (or rather an iteration of speculation and observation). [Unless one counts both Galilean relativity and some parts of the theory of electromagnetism as being contained in medieval knowledge.] Perhaps it would b…
Which doesn't really stop you from training an LLM on 1904 knowledge and having it derive special relativity. LLMs trained with knowledge cutoffs that far back is a fun exercise, and one I've also dabbled in in the past. But you don't have anywhere near enough training data to reach SotA levels of intelligence, even if you had the money for those training runs. And lobotomizing all modern information out of LLMs doesn't seem viable either. So I don't see how you could ever turn that into a viable benchmark
Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#13I feel like AI development is hitting a wall now. Evaluations are driven by benchmarks, but I have no idea who is even evaluating the validity of these benchmarks. The idea that continuously training and scaling models will automatically yield broad capabilities across multiple domains hardly seems true anymore. Benchmarks like Senior SWE-Bench apply discontinuous and arbitrary thresholds to evaluate results. If the…
People have been saying this for at least 2 years now.
> token prices are skyrocketing
Today's SotA (fable and astra @ 50$ /Mtok output) are cheaper than o1-preview (sept '24, 60$ /Mtok output).
And other models are workhorses, with much better capabilities, are at least 1 oom cheaper today than o1-preview. (I'm using this model, since it was the first "thinking" model)
> it feels impossible for this approach to do something like
The models have just provided lean proofs for FLT (a ~1M$ project that was expected to take a human expert ~5 years to complete) and a Millennium prize problem. These are current, relevant, and previously unsolved problems. The obsession people have with "proving relativity from stone-age data" is just moving the goalposts.
Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#14Earlier quoted context omitted.
Nitpicking: It is not possible to derive the theory of relativity from medieval knowledge. At least not in any way resembling how the development actually happened, that is, heavily influenced by observations (or rather an iteration of speculation and observation). [Unless one counts both Galilean relativity and some parts of the theory of electromagnetism as being contained in medieval knowledge.] Perhaps it would b…
The earliest you could reasonably derive special relativity is probably 1881 with the first version of the Michelson–Morley experiment, which showed that speed of light is the same no matter how you move through space. Which, on the scale of discoveries, is a pretty short timeline to the 1905 publication of special relativity. Which doesn't really stop you from training an LLM on 1904 knowledge and having it derive s…
Thinking otherwise mistakes a fundamental error that the only way for something to happen is how it historically happened. Reconciling observations forced relativity historically, but it could have been sussed out without those. That there were very fast but finite speed phenomena (which could motivate relativity reducing to Newtonian models) was seen in 1676 by Rømer with the speed of light and eclipses of Io, a Galilean moon of Jupiter. It probably could have been done with Galileo's telescopes in 1610. (That is just one example of a phenomenon with a very fast yet finite speed to suggest other experiments to test that "hypothetical relativity".)
TBH, this all relates to teaching "physics without calculus" and that sort of thing. Is the most clear presentation/derivation that which mirrors the history of our muddled yet ever demuddling ideas or that which starts from our most thoroughly demuddled ideas, best notations, etc.? The momentum is certainly the historical approach, yet there are notable exceptions.
Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#15I feel like AI development is hitting a wall now. Evaluations are driven by benchmarks, but I have no idea who is even evaluating the validity of these benchmarks. The idea that continuously training and scaling models will automatically yield broad capabilities across multiple domains hardly seems true anymore. Benchmarks like Senior SWE-Bench apply discontinuous and arbitrary thresholds to evaluate results. If the…
> AI development is hitting a wall now People have been saying this for at least 2 years now. > token prices are skyrocketing Today's SotA (fable and astra @ 50$ /Mtok output) are cheaper than o1-preview (sept '24, 60$ /Mtok output). And other models are workhorses, with much better capabilities, are at least 1 oom cheaper today than o1-preview. (I'm using this model, since it was the first "thinking" model) > it fee…
For each generation of advancement the "AI psychosis" of the previous wave wears off. Those who believed 4.6's reasoning was an accurate account of its behaviour; those who believed it had goals and solved useful problems reliably; and so on -- now, attribute only these things to Fable. And no doubt when Fable 6 comes along, it will be only v. 6 that does that.
We have seen fairly marginal progress in LLM reliability and performance since the meaningful start of the high-inference/high-reasoning harness era. It just takes people a few model version bumps to break out of the addiction loop to realise this.
At some point progress will stall entirely, and a couple years after that the spell will break and people will stop treating LLMs "as AI" in the wide-eyed sense, and start treating them as unreliable tools that map Text->Text -- as they do now with earlier model versions.
Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#16Earlier quoted context omitted.
The earliest you could reasonably derive special relativity is probably 1881 with the first version of the Michelson–Morley experiment, which showed that speed of light is the same no matter how you move through space. Which, on the scale of discoveries, is a pretty short timeline to the 1905 publication of special relativity. Which doesn't really stop you from training an LLM on 1904 knowledge and having it derive s…
This is something of an aside on your first paragraph and lh712's. The theory of special relativity (and field theory in general) can be "derived" (as done in e.g., Landau, Lifshitz _Classical Theory Of Fields_) without empirical observation using concepts known to Galileo/Newton (if they count as "Medieval") - just with different assumptions/ideas about inertial frames. If you assume there is any phenomenon with a f…
(1) The relativity principle is exactly what you refer to yourself. You can't apply "Landau--Lifshitz"-like arguments without it. (And I don't think these principles count as medieval knowledge.)
(2) I mentioned electromagnetism, because you need some clue for the concept of an absolute speed that is same for all inertial observers. This is very counterintuitive from our everyday experience, and counterintuitive from the point of view of somebody living in the time of "Galilean" or Newtonian mechanics. Theory of electromagnetism is the only thing that I am aware of, that is nearly (by a stretch) accessible at the level of "medieval" knowledge, from which the concept of constant speed follows. (Historically: Maxwell's completion of previously inconsistent equations of electromagnetism yielded a wave solution propagating with the constant speed of light. This was interpreted in terms of the ether originally, but it is a strong hint in itself for Einsteinian principle of relativity; and a reasoning along the lines you alluded to can them be applied, at least in principle.)
Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#17Earlier quoted context omitted.
I'm a little bit confused about these tire claims as > down to 0C / 32 F (he didn't test colder conditions) I get that people live in different places, but that's a huge caveat. How's that winter if you're not below 0°C? That sounds like "winter tires are worse than all-seasons tires in winter if you exclude winter".
I agree that testing colder conditions would be useful, but I guess you get into problems with it being ice and snow that you're testing, not the "summer tyres get too hard in the cold" hypothesis. It's winter if it's noticeably cooler than summer. For instance London's climate has a January average minimum temperature of just under 3C, which is very clearly winter compared to the July minimum of 14C. (figures from M…
If it's not actively snowing for a few days roads get clear by thawing during day because of salt/cars being hot/sun shining.
The "summer tyres get too hard in the cold" idea could very well come from places where the typical winter is at least < -5°C. It's not uncommon in central/eastern/northern europe for temperatures to be < -10°C for extended periods of time.
Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#18(There's https://withspecific.com/benchmarks/real-swe but since the tasks are private we still don't really know what they're representative of.)
Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#19I feel like AI development is hitting a wall now. Evaluations are driven by benchmarks, but I have no idea who is even evaluating the validity of these benchmarks. The idea that continuously training and scaling models will automatically yield broad capabilities across multiple domains hardly seems true anymore. Benchmarks like Senior SWE-Bench apply discontinuous and arbitrary thresholds to evaluate results. If the…
Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#20Earlier quoted context omitted.
> AI development is hitting a wall now People have been saying this for at least 2 years now. > token prices are skyrocketing Today's SotA (fable and astra @ 50$ /Mtok output) are cheaper than o1-preview (sept '24, 60$ /Mtok output). And other models are workhorses, with much better capabilities, are at least 1 oom cheaper today than o1-preview. (I'm using this model, since it was the first "thinking" model) > it fee…
No one spend ~10 mil USD using gpt 5.1 to solve a millenium prize problem with similar solutions in the training data. So we have no reference for whether 5.1, with that training data and that amount of moeny, could likewise produce such a solution. For each generation of advancement the "AI psychosis" of the previous wave wears off. Those who believed 4.6's reasoning was an accurate account of its behaviour; those w…
Note also when a new benchmark is released, older models do worse on it than recent models, despite none of them being trained to the benchmark.