Live data from Hacker News

OpenAI's GPT-6 Astra on ARC-AGI-3

arcprize.org

131–140 of 151 posts

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#131

Now that ARC-AGI-3 is saturated, with which version number are they going to "certify" that we have reached AGI? Give a number in the replies to this comment and we will check the answers when AGI is here (if so...)

repeating myself, but:

once 3 is solved, we would come up with 4. then 5, 6...

it will be AGI when we cannot come up with a task easy for human but hard for machines. thet's the whole point.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#132

Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated

You mean any repeatable benchmark will be saturated.

The problem is that there is a huge perverse incentive. The intelligence is in the training layer not in the model parameters, but the intelligence is really good at remembering things, so if you let it take the test, it can RL it.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#133

Now that ARC-AGI-3 is saturated, with which version number are they going to "certify" that we have reached AGI? Give a number in the replies to this comment and we will check the answers when AGI is here (if so...)

repeating myself, but: once 3 is solved, we would come up with 4. then 5, 6... it will be AGI when we cannot come up with a task easy for human but hard for machines. thet's the whole point.

exactly, this isn't AGI.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#134

Now that ARC-AGI-3 is saturated, with which version number are they going to "certify" that we have reached AGI? Give a number in the replies to this comment and we will check the answers when AGI is here (if so...)

repeating myself, but: once 3 is solved, we would come up with 4. then 5, 6... it will be AGI when we cannot come up with a task easy for human but hard for machines. thet's the whole point.

[deleted]

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#135
post #54

Earlier quoted context omitted.

It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…

Pure software benchmarks might be getting saturated, but physical ones aren't. Let LLM control a physical robot to perform tasks that average human can do.

Not an llm, but: https://youtu.be/SzvvhRPj6eU?is=4ztgwim0GmjE1h-Y

That (or a near future one) combined with an llm would be insane

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#136
post #51

Earlier quoted context omitted.

It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles. I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the pro…

I don't think IQ is a good measure for intelligence at all. Neither dolphins or octopuses can solve IQ tests.

That sounds a bit like saying that the words-in-noise test is not a good hearing test because neither dolphins nor octopuses can speak.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#137
post #74
post #65

Earlier quoted context omitted.

> Because the IQ is No, it is just because they have difficulties at the bench. > how gives we don't measure them by We'd measure them by all the tests available. Not all test are usable in all circumstances.

> No, it is just because they have difficulties at the bench. I'll put it in another way. A "gifted kid" can be measured incredibly well on an IQ test, but fail miserably at incredibly normal but very difficult tasks such as consoling someone for their loss and managing family crisis. This is a clear example where an IQ measure doesn't translate to a person being capable of meaningfully changing their environments fo…

I don’t follow your logic. “The gifted person is highly intelligent, just not good at deadlifting 500kg” does not diminish the 500kg deadlift and I wouldn’t expect an IQ test to measure it.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#138
post #104

Earlier quoted context omitted.

So vague it’s useless. We cannot in a declarative sense define what is economically valuable work even now let alone into the future. People take what they can get for pay. Very few individuals can demand a wage. The value of employment is obviously designed around that, not some arbitrary definition of “valuable”. Of course an AI will accept $0/hr, it doesn’t mean it does the job. Anyone who could accurately define…

That's the challenge of defining intelligence. What is intelligence? Do you want to take a stab at defining/quantifying it? I'm s afraid anything specific you can come up with will also be useless.

No it isn’t. Nobody said you had to equate “intelligence” with “value of work” and you got it wrong:

I said it’s impossible to define that value, not intelligence. I’ve seen a dozen or so useful definitions of intelligence.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#139
post #95
post #93

Earlier quoted context omitted.

> Knowledge and thus intelligence How can you conflate the two. > solving consciousness We are very much not interested in that. We just need a proper problem solver.

Right, you are interested in “AGI” and presuming none of that requires consciousness right? For example, how do you know that “feeling pain” is not a functional prerequisite for a task. And that consciousness is a prerequisite for feeling pain

Because there is no feeling of pain involved during the reasoning about "How to improve the balance of power in the Indo-Pacific" for the human reasoner, hence there is no need to have any experience of pain for any reasoner.

Re: OpenAI's GPT-6 Astra on ARC-AGI-3

#140
post #93

Earlier quoted context omitted.

> Knowledge and thus intelligence How can you conflate the two. > solving consciousness We are very much not interested in that. We just need a proper problem solver.

Unproven, but it could very well be some problems require consciousness to be solved.

Extraordinary claims require the effort of the proponent so that they can be accepted on the table and take some proportionate share of it.
Post reply on HN