Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
OpenAI's GPT-6 Astra on ARC-AGI-3
21–30 of 129 posts
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#22What are these numbers? Why do they add up to a few hundred thousand dollars? Who paid for that? With what?
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#23Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#24Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#25Earlier quoted context omitted.
There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!
I am no where close to qualified to do that. Hell, experts can't even define what intelligence is, much less define a test for it
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#2699.9% with the right harness? Ok, we're at AGI then. Prediction: We will now see the goalposts moved towards "well, a human costs less / is more efficient" - that will prevail for a few months until they come up with some other test that humans can do easily but is hard for the bots. This cycle will continue for ever and in 25 years, despite having hyper intelligent embodied robots or whatever, we'll still be arguing…
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#27Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#28What are these numbers? Why do they add up to a few hundred thousand dollars? Who paid for that? With what?
OpenAI provides API key with ~unlimited use?
Astra please create a benchmark that’s favorable to your reasoning skills with a human interface but don’t make the score too attainable add some small issues that keep you below 100% to look sensible and to keep my evaluation metric side gig going.
Alignment++
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#29Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#30what the hell is that score/cost curve lol
DeepSeek v4 Flash recently had a similar "more reasoning is cheaper" curve. It's a fun counterintuition.
The question I have is how far back that curve can go without relying on economies of scale to just drag all the points back to the left. And without overfitting a specific metric that I don’t need (like this test).