Live data from Hacker News

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

arxiv.org

121–130 of 139 posts

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#121
post #8

Earlier quoted context omitted.

I really don’t think we know enough about what „intelligence“ is or how LLMs actually work to confidently say that this is the end of the road for LLM.

I'm sorry, we know exactly how LLMs work, this myth that we "dont know how they work" was perpetuated by executives that dont know how they work. We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces.

> We know exactly how attention layers work and how they produce the next word as well as draw them from larger feature spaces.

This is not what's meant by the statements that we don't know how LLMs work. Explain why LLMs are so good at programming, finding bugs, and developing mathematical proofs. Like, way better than all prior tools specifically designed to be bug finding tools, despite being merely "language models".

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#122

This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…

I think LLMs will continue to improve in the capabilities which they are demonstrably good at, but there are many things which they are not good at which it is not cost effective or meaningful to improve, and in these areas we will not consider them “intelligent,” in the same way that we don’t consider computers “intelligent” but do find them very good at doing wrote calculations.

> but there are many things which they are not good at which it is not cost effective or meaningful to improve

Can you name a few such things so I can keep an eye on them in the coming years?

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#123

Earlier quoted context omitted.

I think LLMs will continue to improve in the capabilities which they are demonstrably good at, but there are many things which they are not good at which it is not cost effective or meaningful to improve, and in these areas we will not consider them “intelligent,” in the same way that we don’t consider computers “intelligent” but do find them very good at doing wrote calculations.

> but there are many things which they are not good at which it is not cost effective or meaningful to improve Can you name a few such things so I can keep an eye on them in the coming years?

Not explicitly, no, but its pretty obvious if you work in the industry and aren't blinded by AGI hype

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#124

Earlier quoted context omitted.

We know plenty about human cognition, and we know everything about how LLMs work. True we don’t know anything about intelligence but that is because “intelligence” is it self a fraught and vague term, and we haven’t (and perhaps never will) settled on what it means exactly.

> we know everything about how LLMs work No we don't. That we understand the low level mechanics of a system doesn't mean we understand how any high level phenomena emerge from those low level mechanics. This is as true for quantum mechanics as it is for LLMs.

But we do, and by your logic we can pick any applied statistics, say a Bayesian inference model, and claim we don‘t understand its emergent properties. Heck, we can take any sort of applied mathematics or science, say linguistics, and claim we don‘t understand the emergent properties of language (legal analysis) even though we understand its syntax and phonology.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#125
post #99

Earlier quoted context omitted.

I wish we knew everything about how LLMs work! We only know very basic elements related to their construction and traning dynamics, and pretty much every major question we would like to address still has unknown or vague heuristic answers. This is expected for such a young field of study. In physics, we know the Schrodinger equation, but we dont know everything about how the world works or how to create new materials…

Both you and your sibling are approaching LLMs like it is some sort of science. If you do that there is no wonder you have a lot of unanswered questions. LLMs are not a science, they are applied statistics. Making predictions to evaluate hypothesis and constructing theories around the hyperparameters of LLMs is no different then making predictions to evaluate hypothesis and constructing theories around the configurat…

My point is exactly that we cannot even begin to start the process that will lead to “know everything” there is to know unless we make it a science first. Ad hoc statistical models are different than understanding or complete knowledge of a subject.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#126

Earlier quoted context omitted.

Why would that be the default assumption?

Because they’ve gotten good enough at lots of other things, and the RL keeps improving them, so enough RL should make them good enough at the focus areas.

I've gotten pretty strong in the gym, my bench has improved to two plates. I see no reason why it won't continue to improve until I can bench my house.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#127

Earlier quoted context omitted.

> but there are many things which they are not good at which it is not cost effective or meaningful to improve Can you name a few such things so I can keep an eye on them in the coming years?

Not explicitly, no, but its pretty obvious if you work in the industry and aren't blinded by AGI hype

That's just vibes then. How is that supposed to be convincing?

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#128

Earlier quoted context omitted.

> we know everything about how LLMs work No we don't. That we understand the low level mechanics of a system doesn't mean we understand how any high level phenomena emerge from those low level mechanics. This is as true for quantum mechanics as it is for LLMs.

But we do, and by your logic we can pick any applied statistics, say a Bayesian inference model, and claim we don‘t understand its emergent properties. Heck, we can take any sort of applied mathematics or science, say linguistics, and claim we don‘t understand the emergent properties of language (legal analysis) even though we understand its syntax and phonology.

Yes. That's why law is a different discipline from linguistics requiring its own rules of analysis, pedagogy, etc.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#129

Earlier quoted context omitted.

Please imagine that the benchmark runners are not making mistakes of the fourth grade level - which they won't be. If you assume this level of incompetence, then literally everything is possible.

Why would I assume anything other than maximum incompetence from the AI ecosystem?

This is such a ridiculous and shallow cop-out.

And it also is completely irrelevant to my challenge to show how proprietary benchmarking can be gamed, because it presumes (absolutely insane and divorced from reality) circumstances that have nothing to do with benchmarking as a concept or process.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#130

Earlier quoted context omitted.

Not explicitly, no, but its pretty obvious if you work in the industry and aren't blinded by AGI hype

That's just vibes then. How is that supposed to be convincing?

Unless you are working for a competing firm or have a specific clause in your contract, you can literally just sign up for AI projects as a contractor and start contributing. You will see and understand everything after a few months. Nobody is hiding this knowledge really.
Post reply on HN