Live data from Hacker News

ARC-AGI Leaderboard

arcprize.org

131–140 of 156 posts

Re: ARC-AGI Leaderboard

#131

Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883

And being worst than previous model...

"...The traces tell the why: (1) On our most classic Witness-style game, Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration. It already knows this genre. (2) But on our most novel game (unusual mechanic combinations you can't pattern-match), Opus 5 regresses below Opus 4.8. Where rules must actually be discovered through interaction, the new model is worse than the old one..."

Re: ARC-AGI Leaderboard

#132

Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883

>That decomposition (perfect on templates, regressed on novelty) is the signature of “scaffold-then-internalize” training on genre-specific data, not a general gain in interactive abstract reasoning.

They're smuggling a claim that benchmarks like ARC-AGI measure "interactive abstract reasoning" here, which is what is claimed by the people that make these benchmarks, and also not proven.

Re: ARC-AGI Leaderboard

#133
post #31

Earlier quoted context omitted.

I don't know why exactly, but Fable has felt the most human LLM to arrive.

I wrote this in June, and I'm honestly not sure I've felt the same magic since: I was close to maxing out my $200 plan for the week, almost all Fable use [Claude CLI]. My observations: Fable seemed to have bigger-picture thinking and completed tasks more thoroughly vs just focusing on executing the ask. It pieced together context and intent like an all-star employee would, vs one that just does what you say. Not over…

This is exactly my experience as well.

Re: ARC-AGI Leaderboard

#134
post #130

Earlier quoted context omitted.

You really think they saw the jump in arc-agi-3 (which they reported in their official card), and didn't even bother to check? They maybe have not intentionally benchmaxxed, but they certainly know that's what happened .

How do you propose they check for something like this? They can't exactly ctrl-f the model weights for "Arc-AGI".

> they can’t possibly know or find out what was in the training data

doesn’t appear to be a very strong argument

Re: ARC-AGI Leaderboard

#135
post #62

The last time I checked, for the arc agi 3 leaderboard, the models are given a simple prompt and the game input and asked to play the game, no harness/tools. If harnesses were allowed, I would expect the benchmark to be saturated. There were a few harness attempts, but they could only be evaluated on the public set, so it's not an apples to apples comparison. My guess is, the large score jump for Opus 5 is mainly bec…

The exclusion of harness's feels really weird given that companies are recognizing the value of what harness's can do. By excluding them the benchmark is becoming less relevant.

The goal is to test for AGI where G stands for general, that means ability to act in any environments, ideally solving novel tasks using novel tools we’ve never seen before in the world. If a specific prompt or tool design lifts a model’s score it’s a sign the model is overfitting to a particular modus operandi, therefore not general.

I think in this age where models are heavily RL-ed on acting in specific harnesses, this type of benchmark is more important than ever, to make sure they’re not in fact moving further away from general intelligence.

Re: ARC-AGI Leaderboard

#136
post #130

Earlier quoted context omitted.

You really think they saw the jump in arc-agi-3 (which they reported in their official card), and didn't even bother to check? They maybe have not intentionally benchmaxxed, but they certainly know that's what happened .

How do you propose they check for something like this? They can't exactly ctrl-f the model weights for "Arc-AGI".

[dead]

Re: ARC-AGI Leaderboard

#137

Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883

"Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration." I guess there is no way this can happen without benchmark being part of the training data??

They state the puzzle is "Witness-like" which I assume means that it follows the rules from the well-known puzzle game "The Witness" which Opus definitely knows.

Re: ARC-AGI Leaderboard

#138
post #122

Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883

I have been claiming that I don't think Chinese AI companies are benchmaxxing harder than American AI companies, which has gotten mixed reception: sometimes people agree, sometimes they disagree. It seems I was wrong. American AI companies might actually be benchmaxxing harder.

We run an evaluation that is designed to be less vulnerable to benchmaxxing because there aren't correct solutions; agents are interacting in the same environment as other agents. And it's private, and our public benchmark is not well known enough for anyone to probably care to benchmax us yet. So I think it's pretty indicative of true relative aptitude.

All models have probably memorized significant swaths of solution sets for popular benchmarks at this point, either accidentally or intentionally, so it's all relative at this point. However, in our experience, Chinese models do benchmax harder. This is also consistent with interacting with Chinese labs soliciting data/environments, who literally asked us for datasets and tasks modeled around and formatted like popular benchmarks.

Opus 5 will be uploaded tomorrow, but we already have the tests locally and it is truly as capable as Fable, but at 81% of the real cost. (And from subjective usage, it has a very different personality)

Data at https://gertlabs.com/rankings

Re: ARC-AGI Leaderboard

#139

Earlier quoted context omitted.

It's like we've come full circle: First people practiced L33t3cod3 problems for interviews Then people built AIs to build software And now the AIs are studying L33t3cod3 problems

Why the 3 rather than e?

That’s the original form of the slang/jargon term that the site’s name derived from.

https://en.wikipedia.org/wiki/Leet

Re: ARC-AGI Leaderboard

#140
post #130

Earlier quoted context omitted.

You really think they saw the jump in arc-agi-3 (which they reported in their official card), and didn't even bother to check? They maybe have not intentionally benchmaxxed, but they certainly know that's what happened .

How do you propose they check for something like this? They can't exactly ctrl-f the model weights for "Arc-AGI".

Anthropic expends tons of compute and effort on understanding internal model states [1]; this kind of thing is right up their alley.

[1]: Recent example: https://www.anthropic.com/research/global-workspace

Post reply on HN