Live data from Hacker News

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

arxiv.org

71–80 of 139 posts

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#71
post #24
post #8

Earlier quoted context omitted.

I really don’t think we know enough about what „intelligence“ is or how LLMs actually work to confidently say that this is the end of the road for LLM.

You aren't contradicting the person.

They definitely are - the OP claimed that we are reaching "the end of the road for LLMs", based on absolutely no data and some handwaving on pareto distribution.

We absolutely don't know enough about LLMs and intelligence to make such a bold (and ridiculous) claim. If anything, all evidence point to the contrary, with new scientific breakthrough achieved across a variety of fields via LLMs.

I've been really struggling to understand how the HN community can so boldly claim that LLMs are going to stop improving or not really smart. I just read it as the "denial" stage of the stages of grief that a good portion of this community is in right now (which is understandable).

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#72

This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…

Even without getting better trained models and only speed increase, the output would be dramatically better. An LLM or non llms that is a billion times faster than now would be so insanely strong in many areas.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#73

Earlier quoted context omitted.

If you know all this, can you explain how these models produce advanced mathematical proofs? (as recently done by OpenAI, for example) I tried to generate the next word to the best of my ability, starting with a mathematical problem, but I did not create a valid proof. How do these LLMs work when they create math proofs to problems not yet solved?

Yes, it turns out that matrix math over a feature space of math works pretty well because unlike poetry or real world work, maths are internally coherent and entirely theoretical.

Funny how you said "it turns out" when the whole point is that we don't understand LLMs - we just empirically see what they are good at.

Claiming that you understand LLMs is similar to saying that you understand how our biology work because you understand evolution. No - you understand the mechanism behind evolution, but not the complexity it produces.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#74
post #46

Earlier quoted context omitted.

Can you quantify this rate of progress? Because someone always comes around and says there's been exponential progress in the past every time someone complains that models just aren't very good. Both can't be true

one of these claimed is backed by data, one is backed by anecdotes, you can decide which you trust more

So far, neither is backed by data.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#75
post #64

Earlier quoted context omitted.

Idk about end of the road, I’m sure they can squeeze out some more performance by curating even more data and doing even more RL. But I would bet that pretty much all of the improvement we’ve seen over the last year with coding has come from RL, not from the models becoming particularly stronger. And this makes sense, if models grow sublinearly with compute. And it seems like they do.

It seems pretty obvious from the steep 'intelligence' drop-off on out-of-distribution tasks that the performance improvement is from throwing untold tens of billions at RL. There are legions of highly skilled people employed solely to feed the RL loop. Evidently effective, but there's an unmistakable feeling this won't ultimately be the way forward.

Well, why not? Won’t it get “good enough” at every task eventually?

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#77

This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…

I think LLMs will continue to improve in the capabilities which they are demonstrably good at, but there are many things which they are not good at which it is not cost effective or meaningful to improve, and in these areas we will not consider them “intelligent,” in the same way that we don’t consider computers “intelligent” but do find them very good at doing wrote calculations.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#78
post #47

Earlier quoted context omitted.

Weird moment for this take. We're seeing some of the fastest and most impressive progress ever right now. Frontier labs have categorically different & better set ups for evaluation, they're fine. It's work but it's not a crisis.

I've heard that every week of every month for the past three years. And yet, ask an LLM about a seahorse emoji and see what happens.

And it held true every week of every month for the past three years. AI progress is screaming forward at a breakneck pace.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#79
post #60
post #46

Earlier quoted context omitted.

Can you quantify this rate of progress? Because someone always comes around and says there's been exponential progress in the past every time someone complains that models just aren't very good. Both can't be true

Both can be true, because the experience depends on the skill of the user. The article the other day here on HN that LLMs reward skill is my exact experience. If you are are good at what you are trying to use it for they can be a skill amplifier, and they are definitely getting much better rapidly for the work I do with them. At the same time people are complaining that they are getting dumber. Saying that both can't…

"Skill" with an LLM is nonsense. It's non-deterministic, so any equal efforts are not to be given equal outcome.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#80
post #60

Earlier quoted context omitted.

Both can be true, because the experience depends on the skill of the user. The article the other day here on HN that LLMs reward skill is my exact experience. If you are are good at what you are trying to use it for they can be a skill amplifier, and they are definitely getting much better rapidly for the work I do with them. At the same time people are complaining that they are getting dumber. Saying that both can't…

"Skill" with an LLM is nonsense. It's non-deterministic, so any equal efforts are not to be given equal outcome.

not equal, but probabisticly better. Best example: give the agent a tight spec and it will perform better compared with a spec that leaves room for interpretation. This is true for all models, more or less. (purely anecdotal of course)
Post reply on HN