Live data from Hacker News

ARC-AGI Leaderboard

arcprize.org

141–150 of 156 posts

Re: ARC-AGI Leaderboard

#141
post #122

Earlier quoted context omitted.

I have been claiming that I don't think Chinese AI companies are benchmaxxing harder than American AI companies, which has gotten mixed reception: sometimes people agree, sometimes they disagree. It seems I was wrong. American AI companies might actually be benchmaxxing harder.

We run an evaluation that is designed to be less vulnerable to benchmaxxing because there aren't correct solutions; agents are interacting in the same environment as other agents. And it's private, and our public benchmark is not well known enough for anyone to probably care to benchmax us yet. So I think it's pretty indicative of true relative aptitude. All models have probably memorized significant swaths of soluti…

Benchmaxxing via memorization is boring and doesn't fool anyone for too long. It works, but then new benchmarks test old models and the real results fall in line. Benchmaxxing by focusing on specific types of things that benchmarks test on, while still not improving intelligence or capability in the general case? Not only is it blatantly obvious that all AI labs do this, but it's not even obvious how you would go about it any other way.

Now I am not really specifically accusing Anthropic of anything here, I'm just saying their behavior is suspicious. Since you tested Fable, they wouldn't even have to lie to have optimized for your specific benchmarks, since they absolutely had permission to read your sessions if they wanted to. But obviously, that's only the situation if we take them at their word. Personally I would be a bit surprised if they just flat out were lying and secretly retaining data they say they are not, but not that surprised. The penalties for doing this are probably worth the rewards if it keeps them super far ahead in the benchmarks for years without anyone catching on.

(In actuality though, even if they really were trying to sneakily grab samples of benchmark tests via their Fable data retention rules, I don't really suspect there would've been very much time to optimize Opus 5 on it. So consider me bothered.)

Re: ARC-AGI Leaderboard

#142

Earlier quoted context omitted.

Meh, I doubt it was intentional. Deliberate benchmaxxing is incredibly damaging to credibility once it's discovered (see what happened to Meta with LlaMa 4). It's more likely that the training data was contaminated with the benchmark data.

You really think they saw the jump in arc-agi-3 (which they reported in their official card), and didn't even bother to check? They maybe have not intentionally benchmaxxed, but they certainly know that's what happened .

Is it really all that different from how the models "know" anything else?

The best way to know that christmas trees are typically purchased in december isn't to access historical financial tables and run the entire analysis yourself but to just "know it" from the articles you read attesting to the fact.

If you've found data that is the answer to the question, be it a random user question or a benchmark, the softwares' goal is to produce the correct solution and it is cheaper to retrieve from storage than to compute.

It's the same with coding. The agent usually isn't really thinking about the problem from first principles, it's just giving you the answers it's already found in instances where someone else asked the same question.

It seems to me that if you're really testing for reasoning capability, just as if you were an instructor administering a test, you'll need to change the test from run to run in order to make sure the agent/student isn't just copying old tests.

Re: ARC-AGI Leaderboard

#143
ARC-AGI3 doesn't seem like a great benchmark to me in the first place. It assumes a lot of human like tendencies which an AI either shouldn't or wouldn't have. Particularly in the genre of "gameplay" where unspoken assumptions from prior games inform our understanding of rules.

Re: ARC-AGI Leaderboard

#144
When we started talking about AGI a few years ago there seemed to be a relatively common consensus that LLM models could not be considered AGI because of how they work. Even if there was some changes to train the model on the fly, I still just don’t feel like this is AGI. It’s just a more convincing version of the existing “party trick” we have been doing all along. It convinces us it’s “AGI” the same way current models would convince someone 10 years ago that it was intelligent.

Nobody has any right to take anything I say seriously, because I’m just some random on the internet. But I don’t think true AGI is any closer than about 10-20 years away. That would be to create an actual analog for a human brain.

Re: ARC-AGI Leaderboard

#147

When we started talking about AGI a few years ago there seemed to be a relatively common consensus that LLM models could not be considered AGI because of how they work. Even if there was some changes to train the model on the fly, I still just don’t feel like this is AGI. It’s just a more convincing version of the existing “party trick” we have been doing all along. It convinces us it’s “AGI” the same way current mod…

AGI and ASI are just a convenient myths that the American AI corps use to push for regulatory capture. The difference between the rhetoric from China surrounding artificial general intelligence, and the rhetoric from America, is pretty stark. The Chinese are a lot more grounded and realistic about the whole thing (they almost never talk about ASI, and only talk about AGI in practical terms), compared to the breathy "humanity is doomed but we're building this shit anyway" stuff coming out of Anthropic.

Re: ARC-AGI Leaderboard

#148
post #47
post #39

Earlier quoted context omitted.

> I've worked with these systems for four years now and they have not meaningfully improved in that time frame. Not meaningfully improved?! Four years ago was gpt *3.5*! ChatGPT hadn’t been released!

Yes! Impressive, isn't it? I see how it has improved for some minor points, that the big models can cover more finetuning ground, but my big gripes are still the same - you could do the same back then with multiple models and more targeted finetuning.

The only thing impressive is how wrong you are. LLMs have improved by an absolutely incredible amount in the last 4 years.

Re: ARC-AGI Leaderboard

#149

Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883

"Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration." I guess there is no way this can happen without benchmark being part of the training data??

No, I believe you are misunderstanding the quote. Each “question” in ARC-AGI-3 is a game that has hidden rules that you can understand if you look at the game board long enough. This quote means that Opus 5 is looking at the game board, figuring out the rules, and writing out the rules before it makes a single move. You can do the same thing if you go to the ARC-AGI-3 website and try some of the games.

Re: ARC-AGI Leaderboard

#150
post #122

Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883

I have been claiming that I don't think Chinese AI companies are benchmaxxing harder than American AI companies, which has gotten mixed reception: sometimes people agree, sometimes they disagree. It seems I was wrong. American AI companies might actually be benchmaxxing harder.

someone's promotion depends on benching harder
Post reply on HN