Live data from Hacker News

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

arxiv.org

91–100 of 139 posts

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#91

As a PC gamer who grew up in the 00s, this has been something I’ve tried to warn ardent LLM and model enthusiasts about for quite some time. Benchmarks are handy when they’re new, novel, and constantly changing. The second you let even a single aspect of it stagnate, it becomes a gameable score rather than a useful metric. In PC Gaming, we saw vendors optimize for specific titles, benchmark tools, and scenarios at th…

> Building a new benchmark won’t solve the problem, either. It will if the benchmark is proprietary. If you can't train on it, then it's extremely difficult to game, and if it's hard enough, then it's economically more efficient to just...make the model smarter

We thought the same thing. Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs).

The only solid way to reduce gaming the benchmarks is to ensure they're always, always changing. Enough pelicans on bicycles, start asking for turtles piloting gliders, or rabbits pushing skateboards. Broaden the focus, not narrow it into repetitive measures.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#92
post #74

Earlier quoted context omitted.

one of these claimed is backed by data, one is backed by anecdotes, you can decide which you trust more

So far, neither is backed by data.

I don’t think that’s fair to say, there are two principle sources of data that paint a fairly consistent picture

- one is scaling laws, where we found years ago that pretraining validation loss scales in an almost miraculously predictable way with data volume and compute. There are apparently theoretical bases for this that I don’t quite understand but this property alone is holding at every scale we’ve ever tested. There is not just “one” scaling law but the point is there are scaling laws and they continue to faithfully predict the performance gains we see

- one is benchmarks, which I always point to epoch capability index as a good summary of them in aggregate which makes it nice to plot on one curve the capability improvement over time

To me either one without the other is substantially weaker, the fact that theory and empirical measurements give you a very good scaling law on a more unintuitive quantity (pretraining validation loss) that’s only indirectly related to the downstream performance you care about, benchmarks (in aggregate) are more direct measures of downstream performance but are harder to nail down clean and well motivated “laws” from theory (as far as I can tell). Nevertheless we do in fact see a clear trend that is not slowing.

That doesn’t mean there aren’t a whole host of benchmark problems that don’t impact the numbers involved here (leakage from training data, fundamental flaws in the design, benchmaxxing) but they don’t change the larger story. These problems don’t plausibly explain the clean trends we see.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#93

This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…

>This seems to be the end of the road for LLM's This is an amazingly ignorant thing to say given the current pace of progress.

What is the current pace of progress?

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#94
post #8

Earlier quoted context omitted.

I really don’t think we know enough about what „intelligence“ is or how LLMs actually work to confidently say that this is the end of the road for LLM.

We know plenty about human cognition, and we know everything about how LLMs work. True we don’t know anything about intelligence but that is because “intelligence” is it self a fraught and vague term, and we haven’t (and perhaps never will) settled on what it means exactly.

I’m certainly no expert in the field but to my knowledge a lot of the LLM science is empirical. I’m not aware of a theory that lets us predict what architecture and what number of parameters is needed to solve a particular set of problems.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#95
post #60
post #46

Earlier quoted context omitted.

Can you quantify this rate of progress? Because someone always comes around and says there's been exponential progress in the past every time someone complains that models just aren't very good. Both can't be true

Both can be true, because the experience depends on the skill of the user. The article the other day here on HN that LLMs reward skill is my exact experience. If you are are good at what you are trying to use it for they can be a skill amplifier, and they are definitely getting much better rapidly for the work I do with them. At the same time people are complaining that they are getting dumber. Saying that both can't…

> At the same time people are complaining that they are getting dumber.

I think this is due to rapidly rising expectations.

When LLMs first show they can do some new thing, we're excited at first. Then, we quickly start taking it for granted, and get upset whenever the LLM fails.

Just three years ago, LLMs could barely hold a conversation. Now, they're writing entire code bases and solving famous mathematical conjectures, but we still focus on whatever they can't do.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#96

Earlier quoted context omitted.

> Building a new benchmark won’t solve the problem, either. It will if the benchmark is proprietary. If you can't train on it, then it's extremely difficult to game, and if it's hard enough, then it's economically more efficient to just...make the model smarter

We thought the same thing. Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs). The only solid way to reduce gaming the benchmarks is to ensure they're always, always changing. Enough pelicans on bicycles, start asking for turtles piloting gliders, or rabbits pushing skateboards. Broaden t…

> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs).

Please explain how, if I'm OpenAI and I'm making ChatGPT 5.7, and I release it, and Artificial Analysis goes off and runs one of their proprietary benchmarks on it from a random account, how I can optimize for that benchmark.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#97

This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…

>This seems to be the end of the road for LLM's This is an amazingly ignorant thing to say given the current pace of progress.

I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for example, has nosedived as they've gotten more intelligent, which makes the frontier models difficult to use even for things like writing emails.

In that sense, the frontier models are going to quickly blaze past any semblance of usefulness to humans, while every once in a while we get a news drop like "GPT-7 solved some crazy math problem" or "it invented some new awesome drug"; meanwhile what most people will use will be smaller, more human-specialized models, maybe distilled from those frontier models, that take much longer to iterate on because they rely on large amounts of human feedback in the domain they're specialized for. In other words, useful progress will probably slow down and become more linear starting in Q4, bounded by the rate at which the humans paying for it say "yes this is a good react website".

(By the way: I earnestly do categorize "inventing a new drug" as non-useful AI progress, counter-intuitively. The drug industry has more ideas for drugs than they know what to do with; "useful progress" is, after the idea is made, validating that it works in humans and doesn't kill the human, and productionizing it. AI will help with this and does, but I have substantial doubt that we'll ever see the drug pipeline speed up to, like, a year from idea to prescription. That would be useful progress, which unfortunately many AI pilled hypermaxers conveniently forget. The invention of a promising new drug, or the solution to an arcane set theory problem, are cherries that, through the diligent labor of humans and AI, may become useful, but progress is rarely made by the lone intellect having an a-ha moment.)

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#98
post #97

Earlier quoted context omitted.

>This seems to be the end of the road for LLM's This is an amazingly ignorant thing to say given the current pace of progress.

I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for…

You must have different use cases than me.

I find even Deepseek Flash v4 0731 even outperforms Opus for me (at 10x the speed two).

Using Sol, Fable, Kimi 3 and other recent models has been unbelievable for me. I didn’t think we’d get to this level for years.

I’m using them for Ruby, TypeScript and Python. In large existing codebases but also lots of tiny tools.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#99
post #8

Earlier quoted context omitted.

I really don’t think we know enough about what „intelligence“ is or how LLMs actually work to confidently say that this is the end of the road for LLM.

We know plenty about human cognition, and we know everything about how LLMs work. True we don’t know anything about intelligence but that is because “intelligence” is it self a fraught and vague term, and we haven’t (and perhaps never will) settled on what it means exactly.

I wish we knew everything about how LLMs work! We only know very basic elements related to their construction and traning dynamics, and pretty much every major question we would like to address still has unknown or vague heuristic answers. This is expected for such a young field of study. In physics, we know the Schrodinger equation, but we dont know everything about how the world works or how to create new materials even though we know that these materials are composed by atoms and we can simulate small collections of them. In cell biology, we know the sequences that make up the DNA of a cell and we approach the time we can build minimal synthetic cells with pieces we understand, but we only scratch the surface of our level of understanding of how the cells actually work and new discoveries are added every day. In biology at large we still keep finding new types of tubes inside human brains—not sure what you mean by plenty, but we certainly have an extremely limited understanding of human cognition compared to what we might have in 50 years from now. It is not just anout LLMs and intelligence—I would like us to be able to answer practical questions about how LLMs work in order to improve general or specialized LLMs even faster than today. We “know” about scaling in an empirical sense, and it certainly has a long way to go, but it does not feel close to a complete understanding.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#100

Earlier quoted context omitted.

If you know all this, can you explain how these models produce advanced mathematical proofs? (as recently done by OpenAI, for example) I tried to generate the next word to the best of my ability, starting with a mathematical problem, but I did not create a valid proof. How do these LLMs work when they create math proofs to problems not yet solved?

It turns out being orders of magnitude faster at searching with the aid of a strong verifier is a great way to generate proofs

I agree these are central components, but to avoid oversimplification and the mistaken belief that modern LLMs do a lot of search during inference: If it was so simple, the traditional computer algebra systems would have reached similar breakthroughs when deployed at large supercomputer centers. This didnt happen because the search space is huge. You definitely also need a fancy learning algorithm. Although these ingredients would suffice (depending on what the learning algorithm is), you probably also want to learn in the absense of a strong verifier at every step, to allow building a fuzzy/erratic sense of the search space that can lead to planning/intuition and allow distant jumps in a targetted direction.
Post reply on HN