Earlier quoted context omitted.
these just need more compute: - earn $100: easy to verify, hard to learn - increase paid subscriptions in an A/B test: easy to verify, hard to learn but we both know these examples go against the spirit of my point
Perhaps, but I think a bigger problem than lack of compute is the cost of rewards. Games like Chess and Go were solved long before self-driving, partly because it's incredibly cheap to acquire the reward of a bad board game decision, relatively to how expensive it is to acquire the cost of a bad driving decision. With driving, acquiring the reward can cost you $20/hr for human supervisors to generate disengagements,…
OpenAI's GPT-6 Astra on ARC-AGI-3
61–70 of 150 posts
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#62Earlier quoted context omitted.
No, but figuring out that you're playing a snake-like puzzle game at all in an extremely general input domain and then solving it in the least number of moves definitely feels like evidence of intelligence.
You forget the benchmark. The human subjects were told they were being timed. If you believe the lowest time is the primary metric you will absolutely trial and error at speed instead of meticulously plan out your moves to minimize that metric. LLMs are not timed and given that it costs tens of thousands of dollars to run this test they're not optimizing for speed. So you've got a deceptive test, with one metric bein…
Not fully relevant: timing is crucial in all-pass tests, not crucial in pass-or-fail tests. I.e.: first of all, they have to be able to reach the goal, and that is already an achievement. Then - and in parallel - the problem solving must also be optimized for efficiency. But "solving" and "efficiency" are non coincident dimensions.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#63Earlier quoted context omitted.
Didn’t I just see a thing about how actual unemployment is at like 24% a few days ago?
According to that interpretation, ~24% is one of the lowest ever. https://www.lisep.org/tru (I have not gone down the rabbit hole to understand how they achieve that 24% number)
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#64Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
I'm also unclear as to how basic inferential logic puzzles spells out intelligence I think if you summed up measures of intelligence as 'can it do basic symbolic logic in a chain with memory' then yes, you've now achieved the intelligence of an e. coli colony [0], congratulations [0] https://journals.aps.org/prx/abstract/10.1103/PhysRevX.10.03...
It spells out a form of intelligence - some can and some cannot.
Those puzzles are an abstraction of a skill which is thought to be exportable in other domains.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#65Earlier quoted context omitted.
Their input and output interfaces are too different from human's and they're not nearly as smart to take our IQ tests, but both dolphins and octopuses can solve complex puzzles tailored for their environment. Those puzzles are the whole reason scientists know that dolphins and octopuses are more intelligent than other animals.
But we do know they're "intelligent" and also smart in an important capacity. So how gives we don't measure them by IQ? Because the IQ is not a good measure of intelligence or smarts.
No, it is just because they have difficulties at the bench.
> how gives we don't measure them by
We'd measure them by all the tests available. Not all test are usable in all circumstances.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#66Earlier quoted context omitted.
It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…
When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there. It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#67Earlier quoted context omitted.
When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there. It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.
To be fair, Mensa is a Venn diagram between IQ and being a douche.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#68Earlier quoted context omitted.
It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…
When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there. It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.
The people at Mensa aren't a valid sample of people who score high on IQ tests, because there is a such a strong selection effect for people with certain personality traits, such as wanting to join a club based on your IQ.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#69Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#70Earlier quoted context omitted.
When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there. It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.
What does having value or interest to you have to do with intelligence?
(Please forgive the flippant response. I believe it cuts to the core of what the parent was intending.)