We're running out of benchmarks to upper bound AI capabilities
1–10 of 12 posts
Re: We're running out of benchmarks to upper bound AI capabilities
#2These models are ridiculously powerful with a blank slate. It's when they get loaded down with all the necessary (and inevitably unnecessary) context to complete the task that they really start to crumble and fold.
Re: We're running out of benchmarks to upper bound AI capabilities
#3Re: We're running out of benchmarks to upper bound AI capabilities
#4Start front loading the models with 5k, 10k, 50k, 100k tokens of messy quasi related context, and then run the benchmarks. These models are ridiculously powerful with a blank slate. It's when they get loaded down with all the necessary (and inevitably unnecessary) context to complete the task that they really start to crumble and fold.
Re: We're running out of benchmarks to upper bound AI capabilities
#5Re: We're running out of benchmarks to upper bound AI capabilities
#6This is the least true thing ever. All LLMs are terrible at ARC-AGI-3. Every video game can be used as a benchmark. You could rank LLMs on how long they can keep a game of Dwarf Fortress running or how fast they can beat GTA5.
Re: We're running out of benchmarks to upper bound AI capabilities
#7This is the least true thing ever. All LLMs are terrible at ARC-AGI-3. Every video game can be used as a benchmark. You could rank LLMs on how long they can keep a game of Dwarf Fortress running or how fast they can beat GTA5.
We already have specialized AI to play video games