Live data from Hacker News

ARC-AGI Leaderboard

arcprize.org

81–90 of 156 posts

Re: ARC-AGI Leaderboard

#81
post #53
post #48

Earlier quoted context omitted.

Let us be more clear: there is no structural jump, no architectural overcoming of the original fault. (Edit: and on a similar point, structural properties such as having static ntetworks, as opposed to continuously learning and improving architectures (such as us), will reveal that there is still road ahead.)

Maybe more fair then would be: “I've worked with these systems for four years now and while they _have_ meaningfully improved in that time frame, they’re not perfect and remain fundamentally flawed in various ways.” You prompt less. You need not inject search results into the context window yourself, a window much larger than years ago. You get code that’s already been run successfully once instead of finding an obvi…

If you go to a LLM without harness, GP original point in completely right.

LLMS by themselves are still shit at math, they still confuse weird correlation to causation every time (and sometimes in ways even a 9 year old would say "no, that's dumb"), and confuse original parameters very often.

I disagree with " "thinking" token generation being on the correct track just to 'no, wait' on an already correct conclusion", because i think that is an effect of the harness, not the LLMs.

80% off all the improvements since ChatGPT4 are in the harnesses, and the LLMs by themselves, while they improved in areas they already were good at (translation especially) did not fix any of they original issues (object permanence, calculusm correlation).

Just run old models in the playground and get them to play chess (maybe make a small custom harness if you feel like it), then replace it with a frontier model (i don't know if you still have API access without harness on US models, but if you don't try K3), you will see LLMs weaknesses were not at all fixed, even marginally. They're way better and not inducing bugs in the code, so that make them usable since Opus4.5 (anyone using them prior to that either had a greenfield project or like spending hours debugging).

Re: ARC-AGI Leaderboard

#82

ARC-AGI is a beauty contest for pigs where the pig's owners compete to see who can apply the lipstick most convincingly.

Great comparison! We only have to take into account that applying lipstick well bears no consequences, but applying it poorly (i. e. new model tanking the benchmark) could amount to potentially losses of billions of dollars for the pig-breeders (AI labs).

Re: ARC-AGI Leaderboard

#83
post #34

Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)

I've worked with these systems for four years now and they have not meaningfully improved in that time frame. We still have: - statistical correlation between two things will always cause one thing to lead to the other, no matter how much you prompt it to not have that connection (to be expected with a stochastic system) - Math completely fails in longer contexts - "thinking" token generation being on the correct tra…

> I've worked with these systems for four years now and they have not meaningfully improved in that time frame.

That's absolutely insane. Is it some case of anti-AI psychosis?

Re: ARC-AGI Leaderboard

#84
post #40

> Only systems which required less than $10,000 to run are shown. (Notes[1]) Am I lost or are their many models on this ranking (Opus 5 included) that clear this?

Many models are much cheaper through their subscriptions' included usage. That could be what's happening here. Claude gives you something like $5000 of tokens on a $200 plan.

Isn't it ~$3000 per week? Extrapolating from the current limit on Pro plans.

Re: ARC-AGI Leaderboard

#85
post #74
post #67

Earlier quoted context omitted.

I have no idea how you could try and hold this argument without being facetious.

Am I missing something that my original points no longer hold for their products? Did it get meaningfully solved? Are your experiences flawless on that front?

What do you hope to achieve here? Posting absurd and ridiculous things and then continuing like if someone would take you seriously after that. Keep going I guess.

Re: ARC-AGI Leaderboard

#86
post #62

The last time I checked, for the arc agi 3 leaderboard, the models are given a simple prompt and the game input and asked to play the game, no harness/tools. If harnesses were allowed, I would expect the benchmark to be saturated. There were a few harness attempts, but they could only be evaluated on the public set, so it's not an apples to apples comparison. My guess is, the large score jump for Opus 5 is mainly bec…

The exclusion of harness's feels really weird given that companies are recognizing the value of what harness's can do. By excluding them the benchmark is becoming less relevant.

Harness can totally change capability of the model

Re: ARC-AGI Leaderboard

#87
post #70
post #69

Earlier quoted context omitted.

You could argue that if you allowed a harness, and that harness was specific for ARC, then you don’t have AGI, you have something that is definitely not general.

What if the harness is developed by the same model in a prior session?

[deleted]

Re: ARC-AGI Leaderboard

#88
post #81
post #53

Earlier quoted context omitted.

Maybe more fair then would be: “I've worked with these systems for four years now and while they _have_ meaningfully improved in that time frame, they’re not perfect and remain fundamentally flawed in various ways.” You prompt less. You need not inject search results into the context window yourself, a window much larger than years ago. You get code that’s already been run successfully once instead of finding an obvi…

If you go to a LLM without harness, GP original point in completely right. LLMS by themselves are still shit at math, they still confuse weird correlation to causation every time (and sometimes in ways even a 9 year old would say "no, that's dumb"), and confuse original parameters very often. I disagree with " "thinking" token generation being on the correct track just to 'no, wait' on an already correct conclusion",…

I agree, harnesses is where everyone improved the most. Our internal experimental tool can now semi-reliably formulate small programs to assert the correctness of their theories, for example. Context length is still a weird factor that we haven't sensibly solved - if anything, the lesson learned was to keep the context as small as possible and do most of the true "thinking" in the harness and temporary generated code.

Re: ARC-AGI Leaderboard

#89

Earlier quoted context omitted.

Many models are much cheaper through their subscriptions' included usage. That could be what's happening here. Claude gives you something like $5000 of tokens on a $200 plan.

Isn't it ~$3000 per week? Extrapolating from the current limit on Pro plans.

They don't exactly say. I was extrapolating from some sessions usage (my figure was per month, so weeks times ~4).

Re: ARC-AGI Leaderboard

#90
post #62

The last time I checked, for the arc agi 3 leaderboard, the models are given a simple prompt and the game input and asked to play the game, no harness/tools. If harnesses were allowed, I would expect the benchmark to be saturated. There were a few harness attempts, but they could only be evaluated on the public set, so it's not an apples to apples comparison. My guess is, the large score jump for Opus 5 is mainly bec…

The exclusion of harness's feels really weird given that companies are recognizing the value of what harness's can do. By excluding them the benchmark is becoming less relevant.

iirc, a harness isn't allowed, but if the LLM wants to write its own tools to solve things that is allowed.
Post reply on HN