Are these typically the type of tasks that are genuinely worth tens of thousands of dollars?
ARC-AGI Leaderboard
151–156 of 156 posts
Re: ARC-AGI Leaderboard
#152ARC-AGI is a terrible benchmark for testing LLMs because LLMs are not made, trained, or tuned for playing games. They are trained on text to respond well to text based questions and do tasks involving modifying text files. They are not designed for playing games, looking at games, or visual puzzles. Also translating games into text input for the LLM skews the test completely. Imagine trying to get a human to solve vi…
Re: ARC-AGI Leaderboard
#153Earlier quoted context omitted.
OP is delusional or deliberately optuse. I work in the space and stare down these systems 12h/day, and saying the systems haven't meaningfully improved is ludicrous.
OP is largely pissed with what OAI/Antrophic are trying to sell as meaningful improvements and the market-bending money they ask for it. I work in the space and we trained LLM models on conceptual tokens, not language tokens, for example. See Symbolic AI and all the attempts at hybrid models. Also, uh, fame and riches are not really my thing. Middle income is fine. My mistake was speaking up here because I got carele…
Re: ARC-AGI Leaderboard
#154Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883
Re: ARC-AGI Leaderboard
#155Earlier quoted context omitted.
For frontier models, not local. https://schema-harness.github.io/
Yes, saw that. They haven't yet released any code. Until they do, treat it with a huuuge grain of salt. In fact treat any 99% result in ML with a huge grain of salt.
Re: ARC-AGI Leaderboard
#156Earlier quoted context omitted.
Yes, saw that. They haven't yet released any code. Until they do, treat it with a huuuge grain of salt. In fact treat any 99% result in ML with a huge grain of salt.
If you stop and think about the problem it really is quite simple. Just need to build a graph of the game state and then run A* to get to the end.