Live data from Hacker News

ARC-AGI-3

arcprize.org

151–160 of 394 posts

Re: ARC-AGI-3

#151

https://x.com/scaling01 has called out a lot of issues with ARC-AGI-3, some of them (directly copied from tweets, with minimal editing): - Human baseline is "defined as the second-best first-run human by action count". Your "regular people" are people who signed up for puzzle solving and you don't compare the score against a human average but against the second best human solution - The scoring doesn't tell you how m…

Lol basically we're saying AI isn't AI if we utilize the strength of computers (being able to compute). There's no reason why AGI should have to be as "sample efficient" as humans if it can achieve the same result in less time.

Let's say an agent needs to do 10 brain surgeries on a human to remove a tumor and a human doctor can do it in a single surgery. I would prefer the human.

"steps" are important to optimize if they have negative externalities.

Re: ARC-AGI-3

#152

Earlier quoted context omitted.

> Lastly, humans use way less energy to solve these in fewer steps, Not if you count all the energy that was necessary to feed, shelter and keep the the human at his preferred temperature so that he can sit in front of a computer and solve the problem.

ok, but thats the same for bulding a data center. Try again.

Oh and who provided the 'food' for the models?

...

People who write the stuff like the poster above you... are bizzaro. Absolutely bizarro. Did the LLM manfiest itself into existence? Wtf.

Edit, just got confirmation about the bizarro-ness after looking at his youtube.

Re: ARC-AGI-3

#153
post #145

This is a very good estimation of AGI. We give humans and AI the same input and measure the results. Kudos to ARC for creating these games. I really wonder why so many people fight against this. We know that AI is useful, we know that AI is researchful, but we want to know if they are what we vaguely define as intelligence. I’ve read the airplanes don’t use wings, or submarines don’t swim. Yes, but this is is not the…

AGI’s 'general' is the wrong word, I thinkg. Humans aren’t general, we’re jagged. Strong in some areas, weak in others, and already surpassed in many domains. LLM are way past us at languages for instance. Calculators passed us at calculating, etc.

We are jagged, but we can smooth that jaggedness if we choose to do so. LLMs stay jagged.

Re: ARC-AGI-3

#154
My issue with AGI benchmarks is you can never tell if you're measuring actual capability or just how much the training data overlapped with the test.

Re: ARC-AGI-3

#155
post #145

This is a very good estimation of AGI. We give humans and AI the same input and measure the results. Kudos to ARC for creating these games. I really wonder why so many people fight against this. We know that AI is useful, we know that AI is researchful, but we want to know if they are what we vaguely define as intelligence. I’ve read the airplanes don’t use wings, or submarines don’t swim. Yes, but this is is not the…

AGI’s 'general' is the wrong word, I thinkg. Humans aren’t general, we’re jagged. Strong in some areas, weak in others, and already surpassed in many domains. LLM are way past us at languages for instance. Calculators passed us at calculating, etc.

Interesting take.

Just to drive that thought further.

What are you suggesting, should we rename it. To me the fundamental question is this.

Do we still have tasks that humans can do better than AIs?.

I like the question. I think another good test is "make money". There are humans that can generate money from their laptop. I don’t think AI will be net positive.

I’ve tried to create a Polymarket trading bot with Opus 4.6. The ideas were full of logical fallacies and many many mistakes.

But also I’m not sure how they would compare against an average human with no statistics background..

I think it’s really to establish if we by AGI mean better than average human or better than best human..

Re: ARC-AGI-3

#156
post #39

> As long as there is a gap between AI and human learning, we do not have AGI. Back in the 90's, Scientific American had an article on AI - I believe this was around the time Deep Blue beat Kasparov at chess. One AI researcher's quote stood out to me: "It's silly to say airplanes don't fly because they don't flap their wings the way birds do." He was saying this with regards to the Turing test, but I think the sentim…

So…calculators are intelligent? How about accountants that failed arithmetic 101 in high-school, are they intelligent? Generally intelligent?

Re: ARC-AGI-3

#157
post #127

Earlier quoted context omitted.

The human testers were provided with their customary inputs, as were the LLMs. I don't see the issue. I guess it could be interesting to provide alternative versions that made available various representations of the same data. Still, I'd expect any AGI to be capable of ingesting more or less any plaintext representation interchangeably.

The issue is that ARC AGI 3 specifically forbids harnesses that humans get to use.

[deleted]

Re: ARC-AGI-3

#158
post #89

Earlier quoted context omitted.

Well, yes, and would hand even more of an advantage to humans. My point is that designing a test around human advantages seems odd and orthogonal to measuring AGI.

The whole point of AGI is "general" intelligence, and for that intelligence to be broadly useful it needs to exist within the context of a human centric world

General intelligence not owning retinas.

Denying proper eyesight harness is like trying to construct speech-to-text model that makes transcripts from air pressure values measured 16k times per second, while human ear does frequency-power measurement and frequency binning due to it's physical construction.

Re: ARC-AGI-3

#160

Earlier quoted context omitted.

ARC has always had that problem but for this round, the score is just too convoluted to be meaningful. I want to know how well the models can solve the problem. I may want to know how 'efficient' they are, but really I don't care if they're solving it in reasonable clock time and/or cost. I certainly do not want them jumbled into one messy convoluted score. 'Reasoning steps' here is just arbitrary and meaningless. No…

The metric is very similar to cost. It seems odd to justify one and not the other.

Cost has utility in the real world and this doesn't. That's the only reason i would tolerate thinking about cost, and even then, i would never bundle it into the same score as the intelligence, because that's just silly.
Post reply on HN