Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated
Yes, but not necessarily under tight budget constraints.
OpenAI's GPT-6 Astra on ARC-AGI-3
111–120 of 148 posts
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#112Earlier quoted context omitted.
What does IQ have to do with intelligence? (Please forgive the flippant response. I believe it cuts to the core of what the parent was intending.)
Is it a serious question? IQ tries to quantify the positive correlation between the results of all intellectual tasks a person takes (AKA positive manifold).
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#113I like Erdos problems as a benchmark. Models continue to solve them, but a pretty tepid rate now that the low hanging fruit has been taken. From https://epoch.ai/latest/announcing-frontiermath-erdos > Only GPT-6 Astra solved anything: 2 of the 68 problems. It disproved problem 74 by finding a counterexample, at a cost of $218 and 15 hours of working time, and it proved problem 126, at a cost of $247 and 16 hours > Ac…
"Problems solved before a model's training cutoff can be filtered out, and all models compared on the remaining problems" means that the problems an older model actually solved are the ones that get filtered out, while the remaining problems are the ones it already tried and failed on. So older models end up with 0s on the filtered set and you can't really use this to compare new models to older ones.
Also, since these are known public problems, you can't stop people from spending far more than your arbitrary time and $ limits on them. So the number of clean problems will go down over time.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#114I like Erdos problems as a benchmark. Models continue to solve them, but a pretty tepid rate now that the low hanging fruit has been taken. From https://epoch.ai/latest/announcing-frontiermath-erdos > Only GPT-6 Astra solved anything: 2 of the 68 problems. It disproved problem 74 by finding a counterexample, at a cost of $218 and 15 hours of working time, and it proved problem 126, at a cost of $247 and 16 hours > Ac…
A very long tail of problems that weren't solved by humans? Sure. It's a sarcastic take and I understand that you are probably talking about "spiky intelligence", but you've chosen unsolved problems as a measure of the progress yourself.
There are 1217 problems that Erdos proposed, 595 of which are open: https://github.com/teorth/erdosproblems
Unlike traditional benchmarks, it's difficult to overfit your models to produce flattering results to unsolved problems. The open problems very likely do not have published solutions (the initial batches were merely models surfacing data that wasn't published in obvious places, but we're past that now). New models will exhibit something novel by adding solutions. And the matter of solution is interesting as well (contradiction versus a positive proof).
I would venture to predict it'll take years to decades to get to 0 open problems. But I'd be very happy to have this comment look foolish in retrospect as models continue to improve
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#115Earlier quoted context omitted.
When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there. It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.
At the risk of sounding like one of those people at Mensa that annoyed you... The people at Mensa aren't a valid sample of people who score high on IQ tests, because there is a such a strong selection effect for people with certain personality traits, such as wanting to join a club based on your IQ.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#116Earlier quoted context omitted.
Is it a serious question? IQ tries to quantify the positive correlation between the results of all intellectual tasks a person takes (AKA positive manifold).
Whatever the people who came up with IQ intended isn’t really relevant to the question of whether IQ measures intellect. At best is is loosely correlated. Very loosely.
IQ has precise definition. It's what the tests measure. Intellect has only fuzzy handwavey definition. Correlation between IQ and intellect is about as loose as the definition of the intellect. The way people make the correlation even looser is by defining intellect in even more fuzzy and nebulous manner.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#117Earlier quoted context omitted.
A very long tail of problems that weren't solved by humans? Sure. It's a sarcastic take and I understand that you are probably talking about "spiky intelligence", but you've chosen unsolved problems as a measure of the progress yourself.
What does this mean? There are 1217 problems that Erdos proposed, 595 of which are open: https://github.com/teorth/erdosproblems Unlike traditional benchmarks, it's difficult to overfit your models to produce flattering results to unsolved problems. The open problems very likely do not have published solutions (the initial batches were merely models surfacing data that wasn't published in obvious places, but we're pa…
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#118Earlier quoted context omitted.
To be fair, Mensa is a Venn diagram between IQ and being a douche.
I suspect my IQ isn’t quite high enough to join Mensa but I’ve always flirted with the idea of trying to join and getting in just to see what a group of Mensa people are like.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#119Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#120Earlier quoted context omitted.
Is it a serious question? IQ tries to quantify the positive correlation between the results of all intellectual tasks a person takes (AKA positive manifold).
Whatever the people who came up with IQ intended isn’t really relevant to the question of whether IQ measures intellect. At best is is loosely correlated. Very loosely.