This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression. If the Pareto rule is any indicator, 80% of results come from 20% of causes. It seems that we have alot more to learn about intelligence. I am reminded of a statement on truth from an ancient philosopher, this sentiment seems to be exactly the opposite of the LLM training paradigm…
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
101–110 of 139 posts
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#102Earlier quoted context omitted.
>This seems to be the end of the road for LLM's This is an amazingly ignorant thing to say given the current pace of progress.
I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for…
Point being: it’s overall a worse experience even if the model is technically better at a lot of things.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#103Earlier quoted context omitted.
>This seems to be the end of the road for LLM's This is an amazingly ignorant thing to say given the current pace of progress.
I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for…
We are not even close to what AI slowdown looks like.
The whole business side of things are now building Agentic Layer for Business applications. All of this Agentic Layer needs to be build and its happening right now and still needs a little bit of time.
Anthropic and co have the biggest and centralized reinforcement loop on the planet: Millions of people telling them what is good and what not due to thumbs up/down.
And for sure when the businesses are building the agentic layer they might give direct feedback to them.
While in parallel LLMs get better, more generic and a LOT cheaper too.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#104Earlier quoted context omitted.
We thought the same thing. Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs). The only solid way to reduce gaming the benchmarks is to ensure they're always, always changing. Enough pelicans on bicycles, start asking for turtles piloting gliders, or rabbits pushing skateboards. Broaden t…
> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs). Please explain how, if I'm OpenAI and I'm making ChatGPT 5.7, and I release it, and Artificial Analysis goes off and runs one of their proprietary benchmarks on it from a random account, how I can optimize for that benchmark.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#105Earlier quoted context omitted.
I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for…
I agree on Opus. I’ve had more luck with other models. In particular, Opus’s writing style makes one want to… blow their brains out. While it’s not hallucinating too much, and can troubleshoot certain issues extremely well, the comments it leaves are silky smooth and chock full of inscrutable phrases. And it’s a lot slower than it used to be. Point being: it’s overall a worse experience even if the model is technical…
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#106Earlier quoted context omitted.
We thought the same thing. Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs). The only solid way to reduce gaming the benchmarks is to ensure they're always, always changing. Enough pelicans on bicycles, start asking for turtles piloting gliders, or rabbits pushing skateboards. Broaden t…
> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs). Please explain how, if I'm OpenAI and I'm making ChatGPT 5.7, and I release it, and Artificial Analysis goes off and runs one of their proprietary benchmarks on it from a random account, how I can optimize for that benchmark.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#107Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#108Earlier quoted context omitted.
I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for…
"become more linear starting in Q4" We are not even close to what AI slowdown looks like. The whole business side of things are now building Agentic Layer for Business applications. All of this Agentic Layer needs to be build and its happening right now and still needs a little bit of time. Anthropic and co have the biggest and centralized reinforcement loop on the planet: Millions of people telling them what is good…
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#109Earlier quoted context omitted.
"become more linear starting in Q4" We are not even close to what AI slowdown looks like. The whole business side of things are now building Agentic Layer for Business applications. All of this Agentic Layer needs to be build and its happening right now and still needs a little bit of time. Anthropic and co have the biggest and centralized reinforcement loop on the planet: Millions of people telling them what is good…
Cheaper? For whom? As a solo practitioner, I can no longer afford the workloads I was getting for $20/mo in January. Now the same plan being utilized at the same level for the same work hits its limits within a few hours, and runs out of tokens in less than two days.
I mean the token prices in general as certain services were never really using a subscription.
I do run a claude subscripton right now though and since there capacity change, i hit the limit rarely in comparision to the past, but I don't think this will stay as it is.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#110Earlier quoted context omitted.
> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs). Please explain how, if I'm OpenAI and I'm making ChatGPT 5.7, and I release it, and Artificial Analysis goes off and runs one of their proprietary benchmarks on it from a random account, how I can optimize for that benchmark.
Does OpenAI give Artificial Analysis early access to test models? If so it's definitely not "some random account".