Earlier quoted context omitted.
> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs). Please explain how, if I'm OpenAI and I'm making ChatGPT 5.7, and I release it, and Artificial Analysis goes off and runs one of their proprietary benchmarks on it from a random account, how I can optimize for that benchmark.
Search for accounts @artificialanalysis.com. read their chat history. optimise for that.
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
111–120 of 139 posts
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#112Earlier quoted context omitted.
I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for…
You must have different use cases than me. I find even Deepseek Flash v4 0731 even outperforms Opus for me (at 10x the speed two). Using Sol, Fable, Kimi 3 and other recent models has been unbelievable for me. I didn’t think we’d get to this level for years. I’m using them for Ruby, TypeScript and Python. In large existing codebases but also lots of tiny tools.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#113Earlier quoted context omitted.
We know plenty about human cognition, and we know everything about how LLMs work. True we don’t know anything about intelligence but that is because “intelligence” is it self a fraught and vague term, and we haven’t (and perhaps never will) settled on what it means exactly.
I wish we knew everything about how LLMs work! We only know very basic elements related to their construction and traning dynamics, and pretty much every major question we would like to address still has unknown or vague heuristic answers. This is expected for such a young field of study. In physics, we know the Schrodinger equation, but we dont know everything about how the world works or how to create new materials…
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#114Earlier quoted context omitted.
Search for accounts @artificialanalysis.com. read their chat history. optimise for that.
Please imagine that the benchmark runners are not making mistakes of the fourth grade level - which they won't be. If you assume this level of incompetence, then literally everything is possible.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#115Earlier quoted context omitted.
I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for…
"become more linear starting in Q4" We are not even close to what AI slowdown looks like. The whole business side of things are now building Agentic Layer for Business applications. All of this Agentic Layer needs to be build and its happening right now and still needs a little bit of time. Anthropic and co have the biggest and centralized reinforcement loop on the planet: Millions of people telling them what is good…
Critically, you did not quote the most important part of my sentence: "useful progress will probably slow down and become more linear starting in Q4"; your omission of those words is why I believe you don't understand what I'm saying; you didn't find it important to make your point, so you omitted it, when actually it is critical to the entire assertion. You can read my third paragraph, if you wish, to understand why it is important, instead of just stopping at the first word you disagree with and hitting the "Submit Comment" button.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#116Earlier quoted context omitted.
Well, why not? Won’t it get “good enough” at every task eventually?
Why would that be the default assumption?
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#117Earlier quoted context omitted.
Both can be true, because the experience depends on the skill of the user. The article the other day here on HN that LLMs reward skill is my exact experience. If you are are good at what you are trying to use it for they can be a skill amplifier, and they are definitely getting much better rapidly for the work I do with them. At the same time people are complaining that they are getting dumber. Saying that both can't…
"Skill" with an LLM is nonsense. It's non-deterministic, so any equal efforts are not to be given equal outcome.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#118Earlier quoted context omitted.
>This seems to be the end of the road for LLM's This is an amazingly ignorant thing to say given the current pace of progress.
There is high rate of progress in specific domains, not high rate of progress in generalness. The models haven't gotten generally smarter, for things they didn't focus on the models are just as bad as a year ago.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#119Earlier quoted context omitted.
I really don’t think we know enough about what „intelligence“ is or how LLMs actually work to confidently say that this is the end of the road for LLM.
We know plenty about human cognition, and we know everything about how LLMs work. True we don’t know anything about intelligence but that is because “intelligence” is it self a fraught and vague term, and we haven’t (and perhaps never will) settled on what it means exactly.
No we don't. That we understand the low level mechanics of a system doesn't mean we understand how any high level phenomena emerge from those low level mechanics.
This is as true for quantum mechanics as it is for LLMs.
Re: When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
#120Earlier quoted context omitted.
If you know all this, can you explain how these models produce advanced mathematical proofs? (as recently done by OpenAI, for example) I tried to generate the next word to the best of my ability, starting with a mathematical problem, but I did not create a valid proof. How do these LLMs work when they create math proofs to problems not yet solved?
It turns out being orders of magnitude faster at searching with the aid of a strong verifier is a great way to generate proofs