When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
51–60 of 139 posts
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#52I started thinking about this back after the Llama 4 release, and since then our team has put a lot of thought into designing evaluations that don't saturate, are resistant to contamination, and can scale. What has worked best for us is using multi-agent environments with open-ended cooperative or competitive goals. Mostly designed as multiplayer games. The results tend to align with our experience for coding better…
I think it's interesting, I think other people would find it useful, but I don't want to spend a bunch of money running it against all the frontier models.
What's the best way to reach out to labs like yours to collaborate on something like that? Are there any labs that are more open to submissions from internet randos?
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#53Earlier quoted context omitted.
I'm sorry, we know exactly how LLMs work, this myth that we "dont know how they work" was perpetuated by executives that dont know how they work. We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces.
If you know all this, can you explain how these models produce advanced mathematical proofs? (as recently done by OpenAI, for example) I tried to generate the next word to the best of my ability, starting with a mathematical problem, but I did not create a valid proof. How do these LLMs work when they create math proofs to problems not yet solved?
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#54Earlier quoted context omitted.
Benchmarks saturate around 80-90%? This is not "Acing" a test, this is hitting a wall.
Even on very small tests a fraction of questions might have wrong answers in the key. If models can't get more than 90% of the benchmark right I think it's a strong indication that they were not trained on the answers and that benchmark itself is messy enough that <10% desired answers might be wrong or misleading.
Also to respond to the parent comment: benchmarks have a variety of difficulty levels. Humanity’s Last Exam, though now hitting the beginning of a saturation phase with Fable, was long unsaturated while other benchmarks saturated awhile ago. So that’s what I meant by Epoch capability index: using IRT models this effect so that you gather robust signals from variety of benchmark difficulties and can track progress over time as model capabilities have evolved (and so benchmarks have had to evolve to keep up).
But yes like I was saying: all benchmarks are problematic, some are useful. Benchmark quality problems abound, so 90% being the true ceiling is not surprising. There may be other factors at play here too, I haven’t studied this problem that deeply to have a good thorough answer to this. But keep in mind there are probably 50,000 benchmarks in the literature and that is not a joke number. A crapload of noise in that signal but it’s not all noise.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#55Earlier quoted context omitted.
>This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. It's just inadequate benchmarks. Anyone who has used Fable for anything particularly difficult will have seen that it's miles ahead of Opus 5.0, yet the majority of benchmarks are completely unable to capture this.
If most models were getting 100% on the test it would be an inadequate benchmarks. What were seeing is all models failing to ace these tests. "Benchmark Saturation" is term that promotes lowering the bar.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#56Earlier quoted context omitted.
>This seems to be the end of the road for LLM's This is an amazingly ignorant thing to say given the current pace of progress.
Can you quantify this rate of progress? Because someone always comes around and says there's been exponential progress in the past every time someone complains that models just aren't very good. Both can't be true
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#57This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…
>This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. It's just inadequate benchmarks. Anyone who has used Fable for anything particularly difficult will have seen that it's miles ahead of Opus 5.0, yet the majority of benchmarks are completely unable to capture this.
I have seen people claim the exact opposite. If there was so much progress, then you wouldn't have endless disagreements with people championing their own favourite model as the strongest.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#58Earlier quoted context omitted.
I work in AI evaluation, lots of problems and leakage is an issue as is ecological validity, but they definitely do not explain the progress we see. I think Epoch has the best analysis I’ve seen on evaluation trends; they use IRT to basically model a variety of benchmark difficulties, and then model a capability parameter for each model. This is as robust a sort of “meta-study” of evaluations as I’ve seen and the tre…
Benchmarks saturate around 80-90%? This is not "Acing" a test, this is hitting a wall.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#59It appears y combinator has removed this interesting paper from its top trending position. Gee I wonder why?
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#60Earlier quoted context omitted.
>This seems to be the end of the road for LLM's This is an amazingly ignorant thing to say given the current pace of progress.
Can you quantify this rate of progress? Because someone always comes around and says there's been exponential progress in the past every time someone complains that models just aren't very good. Both can't be true
Even if both aren't true, your evidence was people saying two opposing things. The truth (if there is a single objective truth on a given thing) has little bearing on whether or not different people agree on it.