Live data from Hacker News

ARC-AGI Leaderboard

arcprize.org

111–120 of 156 posts

Re: ARC-AGI Leaderboard

#114
post #72

Earlier quoted context omitted.

OP is delusional or deliberately optuse. I work in the space and stare down these systems 12h/day, and saying the systems haven't meaningfully improved is ludicrous.

OP is largely pissed with what OAI/Antrophic are trying to sell as meaningful improvements and the market-bending money they ask for it. I work in the space and we trained LLM models on conceptual tokens, not language tokens, for example. See Symbolic AI and all the attempts at hybrid models. Also, uh, fame and riches are not really my thing. Middle income is fine. My mistake was speaking up here because I got carele…

> are trying to sell as meaningful improvements

There may not be "core" improvements (structural reliability) but there are "emergent" improvements (apparent intelligence). Already the IQ tests from Maxim Lott ( trackingai.org ) show a progressive sliding towards the right side of the curve - which btw translates to a very much non-secondary decline in the user's frustration (and progresses with an increase of usability).

The more they work on it, the more probable the jump becomes - e.g. to achieve the Large Conceptual Models you say you worked on.

Re: ARC-AGI Leaderboard

#115

Earlier quoted context omitted.

It's like we've come full circle: First people practiced L33t3cod3 problems for interviews Then people built AIs to build software And now the AIs are studying L33t3cod3 problems

Why the 3 rather than e?

("Why 'leet' or '1337' instead of 'elite'". Because restricted groups like to stress a difference.)

Re: ARC-AGI Leaderboard

#116
post #72

Earlier quoted context omitted.

OP is delusional or deliberately optuse. I work in the space and stare down these systems 12h/day, and saying the systems haven't meaningfully improved is ludicrous.

OP is largely pissed with what OAI/Antrophic are trying to sell as meaningful improvements and the market-bending money they ask for it. I work in the space and we trained LLM models on conceptual tokens, not language tokens, for example. See Symbolic AI and all the attempts at hybrid models. Also, uh, fame and riches are not really my thing. Middle income is fine. My mistake was speaking up here because I got carele…

You're welcome to speak up, but saying models haven't meaningfully improved in three years, when the most recent models are solving top-level frontier maths problems, is indeed going to strike most people as weird. I still have no idea what you mean.

Re: ARC-AGI Leaderboard

#117
post #77

ARC-AGI-3 launched a few months ago which would suggest that prior models likely had no knowledge of ARC-AGI-3 or training on similar challenges. I could be wrong, but given the large outsized jump solely in the ARC-AGI-3 score, it would suggest that the model didn't become significantly more intelligent overall, but significantly better at solving those specific problems. This could mean one of two things (I think):…

> it would suggest that the model didn't become significantly more intelligent overall, but significantly better at solving those specific problems

We would actually need a test that shows the ability of a model to export its skills to more problems ("interdisciplinarity" etc.).

Re: ARC-AGI Leaderboard

#118
post #116
post #72

Earlier quoted context omitted.

OP is largely pissed with what OAI/Antrophic are trying to sell as meaningful improvements and the market-bending money they ask for it. I work in the space and we trained LLM models on conceptual tokens, not language tokens, for example. See Symbolic AI and all the attempts at hybrid models. Also, uh, fame and riches are not really my thing. Middle income is fine. My mistake was speaking up here because I got carele…

You're welcome to speak up, but saying models haven't meaningfully improved in three years, when the most recent models are solving top-level frontier maths problems, is indeed going to strike most people as weird. I still have no idea what you mean.

I think he means that, while results are there, they are mysteriously emergent, since an analysis of the process reveals it can be faulty.

Re: ARC-AGI Leaderboard

#119
post #81
post #53

Earlier quoted context omitted.

Maybe more fair then would be: “I've worked with these systems for four years now and while they _have_ meaningfully improved in that time frame, they’re not perfect and remain fundamentally flawed in various ways.” You prompt less. You need not inject search results into the context window yourself, a window much larger than years ago. You get code that’s already been run successfully once instead of finding an obvi…

If you go to a LLM without harness, GP original point in completely right. LLMS by themselves are still shit at math, they still confuse weird correlation to causation every time (and sometimes in ways even a 9 year old would say "no, that's dumb"), and confuse original parameters very often. I disagree with " "thinking" token generation being on the correct track just to 'no, wait' on an already correct conclusion",…

Sounds very fair, thanks :)

Re: ARC-AGI Leaderboard

#120
post #73
post #51

Earlier quoted context omitted.

If they use Python to fill the gap, and the end user doesn’t have to know or care, is it unfair to assess this as progress and attribute the progress to the _system_? OK, the core technology that is the language model still can’t math as well as you’d hope, but how about the end result users see from the system when they interface with it? “Did you know humans are better at flying today than they were a thousand year…

You are correct, the frameworks around it have improved. In that regard, my assessment is unfair: I only judge the underlying technology and what is sold by the sota providers, with the premise of what it's like when you start fresh. You can achieve a lot by coding around the issues, but that's kinda against the point of 'AI', is it?

orwin‘s response in a cousin comment helped me see your original valid point on harnessless LLMs!

>You can achieve a lot by coding around the issues, but that's kinda against the point of 'AI', is it?

Will think on that a bit more.

Post reply on HN