Live data from Hacker News

OpenAI's GPT-6 Astra on ARC-AGI-3

arcprize.org

121–130 of 132 posts

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#122
post #114

Earlier quoted context omitted.

What does this mean? There are 1217 problems that Erdos proposed, 595 of which are open: https://github.com/teorth/erdosproblems Unlike traditional benchmarks, it's difficult to overfit your models to produce flattering results to unsolved problems. The open problems very likely do not have published solutions (the initial batches were merely models surfacing data that wasn't published in obvious places, but we're pa…

I mean that it's hard to determine retroactively how much time it would have taken humanity to solve an open problem that was solved by AI. This is a measure that can make ASI look mundane because we don't know how long it would have taken mathematicians to solve a subset of the Erdős problems.

but i feel like its purpose is to measure superintelligence in math. So to be a good measure of it, it can't saturate easily/has to be somewhat mundane at even insanely good levels. (though i do expect that once ais are across all areas/approaches superhuman at math at least 30% will be solved--then probably long-term (like after 2 years) less than 30% will remain unsolved.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#123

Earlier quoted context omitted.

Workforce participation is different than unemployment. Didn't click the link but I suspect that's the case with your 24%

24% sounds correct to me. Many of my friend who completed phd are currently unemployed .... And, some of them started doing random job like uber to make a living. If government don't step up and regulate outsourcing like 100% tax, big problems are coming up.

If you're doing a job you aren't unemployed. Underemployment is measured separately

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#125
post #114

Earlier quoted context omitted.

What does this mean? There are 1217 problems that Erdos proposed, 595 of which are open: https://github.com/teorth/erdosproblems Unlike traditional benchmarks, it's difficult to overfit your models to produce flattering results to unsolved problems. The open problems very likely do not have published solutions (the initial batches were merely models surfacing data that wasn't published in obvious places, but we're pa…

I mean that it's hard to determine retroactively how much time it would have taken humanity to solve an open problem that was solved by AI. This is a measure that can make ASI look mundane because we don't know how long it would have taken mathematicians to solve a subset of the Erdős problems.

Yeah, I wouldn't purport to use this as a measure of intelligence as it applies to humans. I'd leave that to the philosophers, but my intuition is that models are still a long way off of true human-like intelligence (though benefit from certain unfair advantages).

I'm most interested in these problems as a relative measure of performance for successive model generations. If the prior generation couldn't solve a problem but the current one can, that's useful information, especially when we take into account what the proofs look like.

It's not a perfect benchmark, but I prefer it to many others that I see floating around.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#126

Earlier quoted context omitted.

That’s because AGI, like a lot of terms, has no meaning besides what each individual subjectively projects onto it.

I agree, like the average human isn't generally intelligent. IMO, AGI is literally no different from ASI, though people think it is. Like, Imagine you have 1,000 generally intelligent humans working for you (which nobody is really) and you were to point them at your pet project. That would be amazing!

The average human would have average intelligence

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#127
post #17

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!

I think children having learning abilities exceeding LLM test-time learning (currently only happens in-context). But it's unethical to determine the true baseline of a child age 6 spending 6 years learning a radically new skill to mastery--and besides if you apply RL pressure to the AIs it would be able to surpass it. I guess I still believe future AIs should have some form of continual learning at test-time.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#128

Earlier quoted context omitted.

I agree, like the average human isn't generally intelligent. IMO, AGI is literally no different from ASI, though people think it is. Like, Imagine you have 1,000 generally intelligent humans working for you (which nobody is really) and you were to point them at your pet project. That would be amazing!

The average human would have average intelligence

What's the metric on "average human" from a global population of > 8 billion people .. and why would they score mid on a standardised IQ test skewed toward western education / culture?

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#130

$360 per puzzle. When they tested people it took about 10 minutes per puzzle. If price/performance keeps falling at the same rate it has been, this will cost less than US minimum wage humans within two years. Three for Phillipines minimum wage.

Once it figures out a puzzle it could probably be instructed to design a specialized harness for Luna to be able to solve other instances of the same puzzle. Minimum wage workers are not solving novel problems.
Post reply on HN