Live data from Hacker News

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

arxiv.org

101–110 of 139 posts

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#101

This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…

The paper suggests the opposite of your first statement. The benchmarks become useless because the successive models keep saturating them.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#102
post #97

Earlier quoted context omitted.

>This seems to be the end of the road for LLM's This is an amazingly ignorant thing to say given the current pace of progress.

I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for…

I agree on Opus. I’ve had more luck with other models. In particular, Opus’s writing style makes one want to… blow their brains out. While it’s not hallucinating too much, and can troubleshoot certain issues extremely well, the comments it leaves are silky smooth and chock full of inscrutable phrases. And it’s a lot slower than it used to be.

Point being: it’s overall a worse experience even if the model is technically better at a lot of things.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#103
post #97

Earlier quoted context omitted.

>This seems to be the end of the road for LLM's This is an amazingly ignorant thing to say given the current pace of progress.

I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for…

"become more linear starting in Q4"

We are not even close to what AI slowdown looks like.

The whole business side of things are now building Agentic Layer for Business applications. All of this Agentic Layer needs to be build and its happening right now and still needs a little bit of time.

Anthropic and co have the biggest and centralized reinforcement loop on the planet: Millions of people telling them what is good and what not due to thumbs up/down.

And for sure when the businesses are building the agentic layer they might give direct feedback to them.

While in parallel LLMs get better, more generic and a LOT cheaper too.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#104

Earlier quoted context omitted.

We thought the same thing. Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs). The only solid way to reduce gaming the benchmarks is to ensure they're always, always changing. Enough pelicans on bicycles, start asking for turtles piloting gliders, or rabbits pushing skateboards. Broaden t…

> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs). Please explain how, if I'm OpenAI and I'm making ChatGPT 5.7, and I release it, and Artificial Analysis goes off and runs one of their proprietary benchmarks on it from a random account, how I can optimize for that benchmark.

Does OpenAI give Artificial Analysis early access to test models? If so it's definitely not "some random account".

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#105
post #97

Earlier quoted context omitted.

I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for…

I agree on Opus. I’ve had more luck with other models. In particular, Opus’s writing style makes one want to… blow their brains out. While it’s not hallucinating too much, and can troubleshoot certain issues extremely well, the comments it leaves are silky smooth and chock full of inscrutable phrases. And it’s a lot slower than it used to be. Point being: it’s overall a worse experience even if the model is technical…

At some point it started using a lot of jargon instead of just laying it down clearly. Reads a little bit like LinkedInspeak.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#106

Earlier quoted context omitted.

We thought the same thing. Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs). The only solid way to reduce gaming the benchmarks is to ensure they're always, always changing. Enough pelicans on bicycles, start asking for turtles piloting gliders, or rabbits pushing skateboards. Broaden t…

> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs). Please explain how, if I'm OpenAI and I'm making ChatGPT 5.7, and I release it, and Artificial Analysis goes off and runs one of their proprietary benchmarks on it from a random account, how I can optimize for that benchmark.

Search for accounts @artificialanalysis.com. read their chat history. optimise for that.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#108
post #103
post #97

Earlier quoted context omitted.

I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for…

"become more linear starting in Q4" We are not even close to what AI slowdown looks like. The whole business side of things are now building Agentic Layer for Business applications. All of this Agentic Layer needs to be build and its happening right now and still needs a little bit of time. Anthropic and co have the biggest and centralized reinforcement loop on the planet: Millions of people telling them what is good…

Cheaper? For whom? As a solo practitioner, I can no longer afford the workloads I was getting for $20/mo in January. Now the same plan being utilized at the same level for the same work hits its limits within a few hours, and runs out of tokens in less than two days.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#109
post #103

Earlier quoted context omitted.

"become more linear starting in Q4" We are not even close to what AI slowdown looks like. The whole business side of things are now building Agentic Layer for Business applications. All of this Agentic Layer needs to be build and its happening right now and still needs a little bit of time. Anthropic and co have the biggest and centralized reinforcement loop on the planet: Millions of people telling them what is good…

Cheaper? For whom? As a solo practitioner, I can no longer afford the workloads I was getting for $20/mo in January. Now the same plan being utilized at the same level for the same work hits its limits within a few hours, and runs out of tokens in less than two days.

Yes this is unfortunate and not what I meant.

I mean the token prices in general as certain services were never really using a subscription.

I do run a claude subscripton right now though and since there capacity change, i hit the limit rarely in comparision to the past, but I don't think this will stay as it is.

Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

#110

Earlier quoted context omitted.

> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs). Please explain how, if I'm OpenAI and I'm making ChatGPT 5.7, and I release it, and Artificial Analysis goes off and runs one of their proprietary benchmarks on it from a random account, how I can optimize for that benchmark.

Does OpenAI give Artificial Analysis early access to test models? If so it's definitely not "some random account".

Artificial Analysis was an clearly meant to be an example. I was obviously talking about the ideal scenario of proprietary benchmarking, not how it might be being screwed up in practice.
Post reply on HN