This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
11–20 of 139 posts
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#12This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…
It's just inadequate benchmarks. Anyone who has used Fable for anything particularly difficult will have seen that it's miles ahead of Opus 5.0, yet the majority of benchmarks are completely unable to capture this.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#13Earlier quoted context omitted.
Do you draw that conclusion from the fact that AI surprisingly quickly reaches the end of each ruler we try to measure it with?
Can't call it AI like that without discrediting yourself. You mean LLMs?
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#14Earlier quoted context omitted.
Do you draw that conclusion from the fact that AI surprisingly quickly reaches the end of each ruler we try to measure it with?
Can't call it AI like that without discrediting yourself. You mean LLMs?
Most "AI" is really an optimization algorithm in software tools, same as its always been. This really isnt anything new, aside from adding a chatbot / MCP interface to the same tools.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#15Earlier quoted context omitted.
Do you draw that conclusion from the fact that AI surprisingly quickly reaches the end of each ruler we try to measure it with?
Its not really that surprising when models are trained on the exams
I think Epoch has the best analysis I’ve seen on evaluation trends; they use IRT to basically model a variety of benchmark difficulties, and then model a capability parameter for each model. This is as robust a sort of “meta-study” of evaluations as I’ve seen and the trend in capabilities show no sign of slowing down.
So I think people’s feelings clash with reality, and that’s because releases are more frequent and the jumps between releases are smaller, but the growth in capabilities _over time_ has not changed for the better or worse over a very very long period of time.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#16Data at https://gertlabs.com/rankings
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#17Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#18This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…
This is an amazingly ignorant thing to say given the current pace of progress.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#19This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…
Frontier labs have categorically different & better set ups for evaluation, they're fine. It's work but it's not a crisis.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#20This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…
I really don’t think we know enough about what „intelligence“ is or how LLMs actually work to confidently say that this is the end of the road for LLM.
We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces.