Live data from Hacker News

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

arxiv.org

81–90 of 139 posts

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#81
post #60

Earlier quoted context omitted.

Both can be true, because the experience depends on the skill of the user. The article the other day here on HN that LLMs reward skill is my exact experience. If you are are good at what you are trying to use it for they can be a skill amplifier, and they are definitely getting much better rapidly for the work I do with them. At the same time people are complaining that they are getting dumber. Saying that both can't…

"Skill" with an LLM is nonsense. It's non-deterministic, so any equal efforts are not to be given equal outcome.

You could say the same when applied to games with high RNG and chance, such as Slay the Spire 2, and yet those with real skill do perform far better than those without. Those with skill can clear the highest difficulties more often than those with lower skill.

Something being "non-deterministic" is orthogonal to whether or not skill plays a role.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#82

This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…

>This seems to be the end of the road for LLM's This is an amazingly ignorant thing to say given the current pace of progress.

[deleted]

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#83

This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…

>This seems to be the end of the road for LLM's This is an amazingly ignorant thing to say given the current pace of progress.

Shallow dismissals like this one are against HN guidelines because they make for very poor discourse.

We all learned more from the prior comment than we did from this one.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#84

Earlier quoted context omitted.

I'm sorry, we know exactly how LLMs work, this myth that we "dont know how they work" was perpetuated by executives that dont know how they work. We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces.

If you know all this, can you explain how these models produce advanced mathematical proofs? (as recently done by OpenAI, for example) I tried to generate the next word to the best of my ability, starting with a mathematical problem, but I did not create a valid proof. How do these LLMs work when they create math proofs to problems not yet solved?

It turns out being orders of magnitude faster at searching with the aid of a strong verifier is a great way to generate proofs

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#85
post #60

Earlier quoted context omitted.

Both can be true, because the experience depends on the skill of the user. The article the other day here on HN that LLMs reward skill is my exact experience. If you are are good at what you are trying to use it for they can be a skill amplifier, and they are definitely getting much better rapidly for the work I do with them. At the same time people are complaining that they are getting dumber. Saying that both can't…

"Skill" with an LLM is nonsense. It's non-deterministic, so any equal efforts are not to be given equal outcome.

"Skill" at poker is nonsense. It's non-deterministic, so any equal efforts are not to be given equal outcome.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#86
post #64

Earlier quoted context omitted.

It seems pretty obvious from the steep 'intelligence' drop-off on out-of-distribution tasks that the performance improvement is from throwing untold tens of billions at RL. There are legions of highly skilled people employed solely to feed the RL loop. Evidently effective, but there's an unmistakable feeling this won't ultimately be the way forward.

Well, why not? Won’t it get “good enough” at every task eventually?

Why would that be the default assumption?

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#87

Earlier quoted context omitted.

Even on very small tests a fraction of questions might have wrong answers in the key. If models can't get more than 90% of the benchmark right I think it's a strong indication that they were not trained on the answers and that benchmark itself is messy enough that <10% desired answers might be wrong or misleading.

Yea this may explain part of it or all of it, it’s likely a case by case kind of thing. Also to respond to the parent comment: benchmarks have a variety of difficulty levels. Humanity’s Last Exam, though now hitting the beginning of a saturation phase with Fable, was long unsaturated while other benchmarks saturated awhile ago. So that’s what I meant by Epoch capability index: using IRT models this effect so that you…

Saturation is mostly just selection effects in play. Throw out the "90% easiest" of tasks, and what remains is a jagged ladder of high difficulty outliers.

Hard to climb, and hard to measure the climb - because you have less effective data points and the datapoints themselves are less linear, while you're still being subject to the measurement noise.

Not having the mislabeled tasks would reduce the saturation, but it wouldn't drive it to zero. Even without the "infinite difficulty tasks", bell curve would do its thing.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#88
post #73

Earlier quoted context omitted.

Yes, it turns out that matrix math over a feature space of math works pretty well because unlike poetry or real world work, maths are internally coherent and entirely theoretical.

Funny how you said "it turns out" when the whole point is that we don't understand LLMs - we just empirically see what they are good at. Claiming that you understand LLMs is similar to saying that you understand how our biology work because you understand evolution. No - you understand the mechanism behind evolution, but not the complexity it produces.

stop it. This is such reductionist bullshit. By your infinite reductionism definition no one knows how anything works

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#89

Earlier quoted context omitted.

Yes, it turns out that matrix math over a feature space of math works pretty well because unlike poetry or real world work, maths are internally coherent and entirely theoretical.

This not understanding "understanding". > maths are internally coherent and entirely theoretical Nope. This kind of wish-washy thinking is not what we mean by understanding. https://iep.utm.edu/math-inc/

stop with reductionist absurdum.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#90
post #60

Earlier quoted context omitted.

Both can be true, because the experience depends on the skill of the user. The article the other day here on HN that LLMs reward skill is my exact experience. If you are are good at what you are trying to use it for they can be a skill amplifier, and they are definitely getting much better rapidly for the work I do with them. At the same time people are complaining that they are getting dumber. Saying that both can't…

"Skill" with an LLM is nonsense. It's non-deterministic, so any equal efforts are not to be given equal outcome.

Baseball is also non-deterministic, and yet some players are apparently worth a lot more than others.
Post reply on HN