Live data from Hacker News

ARC-AGI-3

arcprize.org

291–300 of 394 posts

Re: ARC-AGI-3

#291

Earlier quoted context omitted.

They don't try to prevent you from deleting them and they don't output anything unless prompted.

"they don't output anything unless prompted" Unprompted they're not unlike a human sleeping or in a coma. Those states don't preclude consciousness in other states.

That's besides the point though.

Re: ARC-AGI-3

#292
post #254

Earlier quoted context omitted.

> However, if it can't figure out to render the json to a visual on its own does it really qualify as AGI? I'd still say the benchmark is doing its job here. Can you render serialized JSON text blob to a visual with your brain only? The model can't do anything better than this - no harness means no tool at all, no way to e.g. implement a visualizer in whatever programming language and run it. Why don't human testers…

Huh. I thought it wasn't supposed to receive any instructions tailored to the task but I didn't understand it to be restricted from accessing truly general tools such as programming languages. To do otherwise is to require pointless hoop jumping as frontier models inevitably get retrained to play games using a json (or other arbitrary) representation at which point it will be natural for them and the real test will b…

This is my understanding as well, I thought tools where allowed.

Re: ARC-AGI-3

#293

Earlier quoted context omitted.

How can you tell that you lack conscious experience and qualia?

They assert that they dont have them, in the same way you (presumably) assert that you do have them. Neither have any further evidence and one is not a prioi more likely than the other.

Yep, this basically. I tend to get along well with solipsists.

Re: ARC-AGI-3

#294
Same question I have for all these benchmarks:

What's going to stop e.g. OpenAI from hiring a bunch of teenagers to play these games non-stop for a month and annotate the game with their logic for deriving the rules, generate a data set based on those playthroughs and fine tuning the next version of chatgpt on all those playthroughs?

Re: ARC-AGI-3

#295

Earlier quoted context omitted.

Not true. We don't have a good definition for intelligence - it's very much an I'll know it when I see it sort of thing. Frontier models are reliably providing high undergraduate to low graduate level customized explanations of highly technical topics at this point. Yet I regularly catch them making errors that a human never would and which betray a fatal lack of any sort of mental model. What are we supposed to make…

I think you are getting caught up on the intelligence part. That is the easy part since AGI doesn't have to be intelligent, it just has to be intelligence. If you look at early chess AI you will see that they are very weak compared to even a beginner human. The level of intelligence does not matter for a chess bot to be considered AI. It is that it is emulating intelligence that makes it AI. >But is it general? I don…

How am I getting caught up on it? I acknowledged that I think frontier models qualify as intelligent but disputed the "general" part. In fact for quite a few years now there have been many non-frontier models that I also consider intelligent within a very narrow domain.

I think stockfish reasonably qualifies as superhuman AI but not even remotely "general". Similarly alphafold.

> Actually solving it is not a requirement for AGI.

I think I see what you're trying to get at but taken as worded that can't possibly be right. Otherwise a dumb-as-a-brick automaton that made an "attempt" to tackle whatever you put in front of it would qualify as AGI.

Re: ARC-AGI-3

#296

Earlier quoted context omitted.

Not true. We don't have a good definition for intelligence - it's very much an I'll know it when I see it sort of thing. Frontier models are reliably providing high undergraduate to low graduate level customized explanations of highly technical topics at this point. Yet I regularly catch them making errors that a human never would and which betray a fatal lack of any sort of mental model. What are we supposed to make…

> Yet I regularly catch them making errors that a human never would I have yet to see a "error" that modern frontier models make that I could not imagine a human making - average humans are way more error prone than the kind of person who posts here thinks, because the social sorting effects of intelligence are so strong you almost never actually interact with people more than a half standard deviation away. (The one…

> you almost never actually interact with people more than a half standard deviation away

I wasn't talking about the average person there but rather those who could also craft the high undergrad to low grad level explanations I referred to.

> This has not been a remotely credible claim for at least the past six months

Well it's happened to me within the past six months (actually within the past month) so I don't know what you want from me. I wasn't claiming that they never exhibit evidence of a mental model (can't prove a negative anyhow). There are cases where they have rendered a detailed explanation to me yet there were issues with it that you simply could not make if you had a working mental model of the subject that matched the level of the explanation provided (IMO obviously). Imagine a toddler spewing a quantum mechanics textbook at you but then uttering something completely absurd that reveals an inherent lack of understanding; not a minor slip up but a fundamental lack of comprehension. Like I said it's really weird and I'm not sure what to make of it nor how to properly articulate the details.

I'm aware it's not a rigorous claim. I have no idea how you'd go about characterizing the phenomenon.

Re: ARC-AGI-3

#297

Same question I have for all these benchmarks: What's going to stop e.g. OpenAI from hiring a bunch of teenagers to play these games non-stop for a month and annotate the game with their logic for deriving the rules, generate a data set based on those playthroughs and fine tuning the next version of chatgpt on all those playthroughs?

Wrong question. I suggest:

1) Do models generalize?

2) If they do, and they generalize from this, is that a win?

Chollet was one of the first “they do not generalize” evangelists. I’d be curious to hear what he thinks now, because a) most disagree with him, and b) this test seems designed to get models that can generalize better at visual long context problem solving and agency, exactly where the bleeding edge is right now for needs with agentic systems.

Re: ARC-AGI-3

#298

Earlier quoted context omitted.

A human can sit down to play a game with unknown rules and write a spec as he goes. If a model can't even figure out to attempt that, let alone succeed at it, then it most certainly isn't an example of "general" intelligence.

> A human can sit down to play a game with unknown rules and write a spec as he goes. Some humans can. Many , if not most humans cannot. A significant enough fraction of humans have trouble putting together Ikea furniture that there are memes about its difficulty. You're vastly overestimating the capabilities of the average human. Working in tech puts you in probably the top ~1-5% of capability to intuit and understa…

Yes, I am aware. However an idealized human can do so. Analogously, there are plenty of humans that can't run an 8 minute mile but if your bipedal robot is physically incapable of ever doing that then it isn't reasonable to claim having achieved human level athletic performance. When it can compete in every Olympic event you can claim human level performance at athletics in general.

If the model can't generalize to arbitrary tasks on its own without any assistance then it doesn't qualify as a general intelligence. AGI to my mind means meeting or exceeding idealized human performance on the vast majority of arbitrary tasks that are cherrypicked to be particularly challenging.

Re: ARC-AGI-3

#299

Earlier quoted context omitted.

AGI’s 'general' is the wrong word, I thinkg. Humans aren’t general, we’re jagged. Strong in some areas, weak in others, and already surpassed in many domains. LLM are way past us at languages for instance. Calculators passed us at calculating, etc.

LLMs haven't passed us in language, a child can learn language with so so much less data than an LLM can

isn't that more like rate of learning? Agreed LLM consume a lot of data.

But your average LLM understands more languages then anyone alive. So super human understanding of various text based languages.

Re: ARC-AGI-3

#300
It's getting pretty old now when Francois Chollet puts out a new ARC challenge, claims definitively that no system is going to crack it without being full blown AGI, the benchmark gets saturated in a few months, he claims the systems definitely aren't AGI then puts out a new challenge that no non AGI system can clear and a few months later.... etc. etc.
Post reply on HN