Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883
ARC-AGI Leaderboard
111–120 of 156 posts
Re: ARC-AGI Leaderboard
#112Re: ARC-AGI Leaderboard
#113Re: ARC-AGI Leaderboard
#114Earlier quoted context omitted.
OP is delusional or deliberately optuse. I work in the space and stare down these systems 12h/day, and saying the systems haven't meaningfully improved is ludicrous.
OP is largely pissed with what OAI/Antrophic are trying to sell as meaningful improvements and the market-bending money they ask for it. I work in the space and we trained LLM models on conceptual tokens, not language tokens, for example. See Symbolic AI and all the attempts at hybrid models. Also, uh, fame and riches are not really my thing. Middle income is fine. My mistake was speaking up here because I got carele…
There may not be "core" improvements (structural reliability) but there are "emergent" improvements (apparent intelligence). Already the IQ tests from Maxim Lott ( trackingai.org ) show a progressive sliding towards the right side of the curve - which btw translates to a very much non-secondary decline in the user's frustration (and progresses with an increase of usability).
The more they work on it, the more probable the jump becomes - e.g. to achieve the Large Conceptual Models you say you worked on.
Re: ARC-AGI Leaderboard
#115Earlier quoted context omitted.
It's like we've come full circle: First people practiced L33t3cod3 problems for interviews Then people built AIs to build software And now the AIs are studying L33t3cod3 problems
Why the 3 rather than e?
Re: ARC-AGI Leaderboard
#116Earlier quoted context omitted.
OP is delusional or deliberately optuse. I work in the space and stare down these systems 12h/day, and saying the systems haven't meaningfully improved is ludicrous.
OP is largely pissed with what OAI/Antrophic are trying to sell as meaningful improvements and the market-bending money they ask for it. I work in the space and we trained LLM models on conceptual tokens, not language tokens, for example. See Symbolic AI and all the attempts at hybrid models. Also, uh, fame and riches are not really my thing. Middle income is fine. My mistake was speaking up here because I got carele…
Re: ARC-AGI Leaderboard
#117ARC-AGI-3 launched a few months ago which would suggest that prior models likely had no knowledge of ARC-AGI-3 or training on similar challenges. I could be wrong, but given the large outsized jump solely in the ARC-AGI-3 score, it would suggest that the model didn't become significantly more intelligent overall, but significantly better at solving those specific problems. This could mean one of two things (I think):…
We would actually need a test that shows the ability of a model to export its skills to more problems ("interdisciplinarity" etc.).
Re: ARC-AGI Leaderboard
#118Earlier quoted context omitted.
OP is largely pissed with what OAI/Antrophic are trying to sell as meaningful improvements and the market-bending money they ask for it. I work in the space and we trained LLM models on conceptual tokens, not language tokens, for example. See Symbolic AI and all the attempts at hybrid models. Also, uh, fame and riches are not really my thing. Middle income is fine. My mistake was speaking up here because I got carele…
You're welcome to speak up, but saying models haven't meaningfully improved in three years, when the most recent models are solving top-level frontier maths problems, is indeed going to strike most people as weird. I still have no idea what you mean.
Re: ARC-AGI Leaderboard
#119Earlier quoted context omitted.
Maybe more fair then would be: “I've worked with these systems for four years now and while they _have_ meaningfully improved in that time frame, they’re not perfect and remain fundamentally flawed in various ways.” You prompt less. You need not inject search results into the context window yourself, a window much larger than years ago. You get code that’s already been run successfully once instead of finding an obvi…
If you go to a LLM without harness, GP original point in completely right. LLMS by themselves are still shit at math, they still confuse weird correlation to causation every time (and sometimes in ways even a 9 year old would say "no, that's dumb"), and confuse original parameters very often. I disagree with " "thinking" token generation being on the correct track just to 'no, wait' on an already correct conclusion",…
Re: ARC-AGI Leaderboard
#120Earlier quoted context omitted.
If they use Python to fill the gap, and the end user doesn’t have to know or care, is it unfair to assess this as progress and attribute the progress to the _system_? OK, the core technology that is the language model still can’t math as well as you’d hope, but how about the end result users see from the system when they interface with it? “Did you know humans are better at flying today than they were a thousand year…
You are correct, the frameworks around it have improved. In that regard, my assessment is unfair: I only judge the underlying technology and what is sold by the sota providers, with the premise of what it's like when you start fresh. You can achieve a lot by coding around the issues, but that's kinda against the point of 'AI', is it?
>You can achieve a lot by coding around the issues, but that's kinda against the point of 'AI', is it?
Will think on that a bit more.