Earlier quoted context omitted.
It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…
Pure software benchmarks might be getting saturated, but physical ones aren't. Let LLM control a physical robot to perform tasks that average human can do.
OpenAI's GPT-6 Astra on ARC-AGI-3
91–100 of 128 posts
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#92Earlier quoted context omitted.
According to that interpretation, ~24% is one of the lowest ever. https://www.lisep.org/tru (I have not gone down the rabbit hole to understand how they achieve that 24% number)
Workforce participation is different than unemployment. Didn't click the link but I suspect that's the case with your 24%
If government don't step up and regulate outsourcing like 100% tax, big problems are coming up.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#93“AGI” never made sense to me. It’s a purely marketing term right? I’ve ignored it thinking it would go away, but it keeps coming up. I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery. Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get…
How can you conflate the two.
> solving consciousness
We are very much not interested in that. We just need a proper problem solver.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#94Earlier quoted context omitted.
It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…
Nit: Turing’s actual imitation game is a party game (like Werewolf/Mafia) and nobody’s even trying to win at that. The LLM’s will just tell you they’re an AI.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#95“AGI” never made sense to me. It’s a purely marketing term right? I’ve ignored it thinking it would go away, but it keeps coming up. I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery. Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get…
> Knowledge and thus intelligence How can you conflate the two. > solving consciousness We are very much not interested in that. We just need a proper problem solver.
For example, how do you know that “feeling pain” is not a functional prerequisite for a task. And that consciousness is a prerequisite for feeling pain
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#96Earlier quoted context omitted.
OpenAI provides API key with ~unlimited use?
So, you’re telling me I need to start a benchmark as a side gig to get a bunch of free compute. Astra please create a benchmark that’s favorable to your reasoning skills with a human interface but don’t make the score too attainable add some small issues that keep you below 100% to look sensible and to keep my evaluation metric side gig going. Alignment++
There is no incentive for OpenAI to subsidize is you if no one reads /reports on your benchmark . They are only going to fund a few that are currently popular .
Community acceptance doesn’t automatically mean the best , it is combination of some level of technical quality and the ability of the promoter to socially influence or get support of influencers .
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#97“AGI” never made sense to me. It’s a purely marketing term right? I’ve ignored it thinking it would go away, but it keeps coming up. I get that consciousness differs from intelligence and that our waking awareness of life is a complete mystery. Knowledge and thus intelligence however I consider as actively being solved by these large ML models. That is, with the right combination of machinery and know-how, you’ll get…
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#98Earlier quoted context omitted.
When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there. It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.
To be fair, Mensa is a Venn diagram between IQ and being a douche.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#99Earlier quoted context omitted.
I'm also unclear as to how basic inferential logic puzzles spells out intelligence I think if you summed up measures of intelligence as 'can it do basic symbolic logic in a chain with memory' then yes, you've now achieved the intelligence of an e. coli colony [0], congratulations [0] https://journals.aps.org/prx/abstract/10.1103/PhysRevX.10.03...
> basic inferential logic puzzles spells out It spells out a form of intelligence - some can and some cannot. Those puzzles are an abstraction of a skill which is thought to be exportable in other domains.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#100Earlier quoted context omitted.
Nit: Turing’s actual imitation game is a party game (like Werewolf/Mafia) and nobody’s even trying to win at that. The LLM’s will just tell you they’re an AI.
This is an artifact of how we deliberately craft these models though. We could just as easily fine tune a model that will believe it is not an AI or will attempt to deceive users asking about it
Also, the skill of the human opponents matters. You'd want to test it against people who have practiced playing the game. Otherwise, it's like the difference between building a chess bot that can win against random undergrads who don't normally play, versus winning against grandmasters. And it's not like there's a pool of skilled human players of the imitation game.