How good are LLMs at doing Mensa tests?
OpenAI's GPT-6 Astra on ARC-AGI-3
121–130 of 142 posts
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#122Earlier quoted context omitted.
What does this mean? There are 1217 problems that Erdos proposed, 595 of which are open: https://github.com/teorth/erdosproblems Unlike traditional benchmarks, it's difficult to overfit your models to produce flattering results to unsolved problems. The open problems very likely do not have published solutions (the initial batches were merely models surfacing data that wasn't published in obvious places, but we're pa…
I mean that it's hard to determine retroactively how much time it would have taken humanity to solve an open problem that was solved by AI. This is a measure that can make ASI look mundane because we don't know how long it would have taken mathematicians to solve a subset of the Erdős problems.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#123Earlier quoted context omitted.
Workforce participation is different than unemployment. Didn't click the link but I suspect that's the case with your 24%
24% sounds correct to me. Many of my friend who completed phd are currently unemployed .... And, some of them started doing random job like uber to make a living. If government don't step up and regulate outsourcing like 100% tax, big problems are coming up.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#124Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#125Earlier quoted context omitted.
What does this mean? There are 1217 problems that Erdos proposed, 595 of which are open: https://github.com/teorth/erdosproblems Unlike traditional benchmarks, it's difficult to overfit your models to produce flattering results to unsolved problems. The open problems very likely do not have published solutions (the initial batches were merely models surfacing data that wasn't published in obvious places, but we're pa…
I mean that it's hard to determine retroactively how much time it would have taken humanity to solve an open problem that was solved by AI. This is a measure that can make ASI look mundane because we don't know how long it would have taken mathematicians to solve a subset of the Erdős problems.
I'm most interested in these problems as a relative measure of performance for successive model generations. If the prior generation couldn't solve a problem but the current one can, that's useful information, especially when we take into account what the proofs look like.
It's not a perfect benchmark, but I prefer it to many others that I see floating around.
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#126Earlier quoted context omitted.
That’s because AGI, like a lot of terms, has no meaning besides what each individual subjectively projects onto it.
I agree, like the average human isn't generally intelligent. IMO, AGI is literally no different from ASI, though people think it is. Like, Imagine you have 1,000 generally intelligent humans working for you (which nobody is really) and you were to point them at your pet project. That would be amazing!
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#127Is solving a snake like puzzle game in the least number of moves really what defines intelligence?
There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#128Earlier quoted context omitted.
I agree, like the average human isn't generally intelligent. IMO, AGI is literally no different from ASI, though people think it is. Like, Imagine you have 1,000 generally intelligent humans working for you (which nobody is really) and you were to point them at your pet project. That would be amazing!
The average human would have average intelligence
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#129Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated
Re: OpenAI's GPT-6 Astra on ARC-AGI-3
#130$360 per puzzle. When they tested people it took about 10 minutes per puzzle. If price/performance keeps falling at the same rate it has been, this will cost less than US minimum wage humans within two years. Three for Phillipines minimum wage.