Live data from Hacker News

ARC-AGI-3

arcprize.org

351–360 of 394 posts

Re: ARC-AGI-3

#351

Earlier quoted context omitted.

> you almost never actually interact with people more than a half standard deviation away I wasn't talking about the average person there but rather those who could also craft the high undergrad to low grad level explanations I referred to. > This has not been a remotely credible claim for at least the past six months Well it's happened to me within the past six months (actually within the past month) so I don't know…

How much of this is expectations setting by the heights models reach? i.e. of we could assess a consistent floor of model performance in a vacuum, would we say it's better at "AGI" than the bottom 0.1% of humans?

Not sure how to answer because we were off on a tangent there about mental models.

I think AGI is two things. Intelligence at a given task, which can be scored versus humans or otherwise. And generalization which is entirely separate. We already have superhuman non-general models in a few domains.

So I don't think that "better than AGI at % of humans" is a sensible statement, at least not initially.

Right now humans generalize to all integers while AI companies keep manually adding additional integers to a finite list and bystanders make claims of generality. If you've still got a finite list you aren't general regardless of how long the list is.

If at some point a model shows up that works on all even integers but not odd ones then I guess you could reasonably claim you had AGI that was 50% of what humans achieve. If a model that generalizes to all the reals shows up then it will have exceeded human generality by an infinite degree. We'll cross those bridges when we come to them - I don't think we're there yet.

Re: ARC-AGI-3

#352

Earlier quoted context omitted.

How much of this is expectations setting by the heights models reach? i.e. of we could assess a consistent floor of model performance in a vacuum, would we say it's better at "AGI" than the bottom 0.1% of humans?

Not sure how to answer because we were off on a tangent there about mental models. I think AGI is two things. Intelligence at a given task, which can be scored versus humans or otherwise. And generalization which is entirely separate. We already have superhuman non-general models in a few domains. So I don't think that "better than AGI at % of humans" is a sensible statement, at least not initially. Right now humans…

Interestingly, I find that the models generalize decently well as long as the "training" (more analogous to that for humans) fits in (small enough) context. That's to say, "in-context learning" seems good enough for real use.

But of course, that's not quite "long term"

Re: ARC-AGI-3

#353
post #213
post #39

> As long as there is a gap between AI and human learning, we do not have AGI. Back in the 90's, Scientific American had an article on AI - I believe this was around the time Deep Blue beat Kasparov at chess. One AI researcher's quote stood out to me: "It's silly to say airplanes don't fly because they don't flap their wings the way birds do." He was saying this with regards to the Turing test, but I think the sentim…

Humans can do a lot of things that don't require intelligence. Artificial intelligence does not need to be 100% human to be AGI.

It needs to pass the most basic concept of learning, which it can’t currently do. Probably wont ever do after listening to dario on his latest podcast run.

Where we are at today is ASI (artificial semi-intelligence). Maybe in 20 years artificial super intelligence can be achieved, but certainly not AGI.

Re: ARC-AGI-3

#354

I was a human tester (I think) for this set of games. I did 25 games in the 90 minutes allotted. IIRC the instructions did mention to minimize action count but the incentives/setup ($5 per game solved) pushed for solve speed over action count. I do recall trying to not just randomly move around while thinking but that was not the primary goal, so I would expect that the baseline for the human solutions have more acti…

I understood minimal actions intuitively, it just made sense? The stamina meter was shrinking with each step, so I recognized it was something to look out for.

Re: ARC-AGI-3

#355

Earlier quoted context omitted.

> If LLMs cannot learn to beat not-that-difficult of games better than young teens, they are not intelligent. I agree, with unresolved questions. Does it count if the LLM writes code which trains a neural network to play the game, and that neural network plays the game better than people do? Does that only count if the LLM tries that solution without a human prompting it to do so?

I disagree that LLMs cannot solve "unsolved problems." This is already happening, and at fundamental mathematical and medical levels (the fields that are the most demanding when it comes to quality). The idea that we haven't taught LLMs to come up with new answers... That doesn't even sound plausible. Just crank up the temperature, and an LLM will throw out so many ideas you'll exhaust yourself trying to sort through…

Compared to AI our brains are order of magnitudes more energy efficient.

Frontier science doesn’t necessarily mean doesn’t mean that its meaningful, there are a bunch of problems that are tedious to solve with existing patterns. Really tedious. So a human prompting it can let an AI solve it, but at the same time that solution might completely not matter at all ever.

I think the real question should be how much does it help when it matters, moneywise.

Like it can create an app, but can it create an app that makes money and somebody cares about?

Like the amounts a human has to intervene to get an app that makes just 10k MRR is probably around in the 1000s, so how we are at really close to AGI?

I’ve tried and promoted the Ralph Loop, but I learnt that the loop just keeps overcomplicating stuff and then you try to simplify the overcomplication and it simplfies the wrong things and enforces the wrong things and in the end you can not move,till a human goes in and entangles it properly.

Re: ARC-AGI-3

#356

Earlier quoted context omitted.

We don't call a calculator intelligent. A calculator is extremely useful, but it is not intelligent. A computer is extremely useful, but it is not intelligent. Airplanes don't have wings, but they're damn sure useful, and also not intelligent. If LLMs cannot learn to beat not-that-difficult of games better than young teens, they are not intelligent. They are extremely useful. But they are not AGI. Words matter.

So your definition of intelligence would be exactly equal to a human or some subset of them you choose? Could a dog solve ARC-AGI? Probably not. I would not say they lack intelligence. Same with a fruit fly. What if the calculator is powered by actual living neurons? I think you need to know where you actually think the difference between organic machine and intelligence is before making blanket statements. A modern…

I don’t doubt he would fool you when your argument is that a fruit fly is not lacking intelligence.

Re: ARC-AGI-3

#357

Earlier quoted context omitted.

Not sure how to answer because we were off on a tangent there about mental models. I think AGI is two things. Intelligence at a given task, which can be scored versus humans or otherwise. And generalization which is entirely separate. We already have superhuman non-general models in a few domains. So I don't think that "better than AGI at % of humans" is a sensible statement, at least not initially. Right now humans…

Interestingly, I find that the models generalize decently well as long as the "training" (more analogous to that for humans) fits in (small enough) context. That's to say, "in-context learning" seems good enough for real use. But of course, that's not quite "long term"

Given that models don't currently learn as they go isn't that exactly what this benchmark is testing? If the model needs to either have been explicitly trained in a similar environment or else to have a human manually input a carefully crafted prompt then it isn't general. The latter case is a human tuning a powerful tool.

If it can add the necessary bits to its own prompt while working on the benchmark then it's generalizing.

Re: ARC-AGI-3

#358

Earlier quoted context omitted.

> Not true. It's certainly true. By definition. If the bar for general intelligence is being smarter than the median human, 50% of people won't reach the threshold for general intelligence. (And if the bar is beating the median in every cognitive test, then a much smaller fraction of people would qualify.) People don't have a consistent definition of AGI, and the definitions have changed over the past couple years, b…

You are using terms like "smart" and "dumb" as if they have universally-accepted definitions. You can make up as many definitions of intelligence as you like (I would argue that is a sign of intelligence) but using those terms is certainly going to lead to circular reasoning.

It has nothing to do with circular reasoning or my personal opinions.

You can choose to define general intelligence in a way that excludes regular people if you like, but then you'd be using a weird definition that differs from how 99.9% of people define it. Humans have general intelligence by any common definition.

Re: ARC-AGI-3

#359

Earlier quoted context omitted.

You are using terms like "smart" and "dumb" as if they have universally-accepted definitions. You can make up as many definitions of intelligence as you like (I would argue that is a sign of intelligence) but using those terms is certainly going to lead to circular reasoning.

It has nothing to do with circular reasoning or my personal opinions. You can choose to define general intelligence in a way that excludes regular people if you like, but then you'd be using a weird definition that differs from how 99.9% of people define it. Humans have general intelligence by any common definition.

Defining it that way doesn't exclude ordinary people. That's an erroneous claim on your part.

Humans as a class exhibit certain capabilities. Thus we expect a class of algorithm to either roughly meet or exceed those capabilities across the board in order to be considered "general". It is clear that we have not yet achieved that.

Re: ARC-AGI-3

#360

Earlier quoted context omitted.

It may have been tested on the full set, but the score you quote is for a single game environment. Not the full public set. That fact is verbatim in what you responded to and vbarrielle quoted. It scored 97% in one game , and 0% in another game. The full prelude to what vbarrielle quoted, the last sentence of which you left out, was: > We then tested the harnesses on the full public set (which researchers did not hav…

> intelligence for those specific games is baked into the harness This is your claim but the other commenter claims the harness consists only of generic tools. What's the reality? I also encountered confusion about this exact issue in another subthread. I had thought that generic tooling was allowed but others believed the benchmark to be limited to ingesting the raw text directly from the API without access to any a…

[deleted]
Post reply on HN