Live data from Hacker News

We're running out of benchmarks to upper bound AI capabilities

lesswrong.com

11–12 of 12 posts

Re: We're running out of benchmarks to upper bound AI capabilities

#11
post #4
post #2

Start front loading the models with 5k, 10k, 50k, 100k tokens of messy quasi related context, and then run the benchmarks. These models are ridiculously powerful with a blank slate. It's when they get loaded down with all the necessary (and inevitably unnecessary) context to complete the task that they really start to crumble and fold.

We need benchmarks that can distinguish between continuous learning and long-context extrapolation.

oh that's easy: continuous learning is not something current architectures can do. So the benchmark for that can be done mentally

Re: We're running out of benchmarks to upper bound AI capabilities

#12

This is the least true thing ever. All LLMs are terrible at ARC-AGI-3. Every video game can be used as a benchmark. You could rank LLMs on how long they can keep a game of Dwarf Fortress running or how fast they can beat GTA5.

ARC-AGI-3 appears to already be saturated [0] [1]

For some reason they refuse to run this on the private set, likely because it's all just a ploy to pump OpenAI

[0] https://arcprize.org/leaderboard/community

[1] https://blog.alexisfox.dev/arcagi3

Post reply on HN