Live data from Hacker News

OpenAI O3 breakthrough high score on ARC-AGI-PUB

arcprize.org

171–180 of 1001 posts

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#171

It sucks that I would love to be excited about this... but I mostly feel anxiety and sadness.

Humanity is about to enter an even steeper hockey stick growth curve. Progressing along the Kardashev scale feels all but inevitable. We will live to see Longevity Escape Velocity. I'm fucking pumped and feel thrilled and excited and proud of our species.

Sure, there will be growing pains, friction, etc. Who cares? There always is with world-changing tech. Always.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#172

[flagged]

> most of the people training these next-gen AIs are neurodiverse Citation needed. This is a huge claim based only on stereotype.

So true. Perhaps I'm just thinking it's my people and need to update my priors.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#173
post #134

Earlier quoted context omitted.

It doesn't need to be general intelligence or perfectly map to human intelligence. All it needs to be is useful. Reading constant comments about LLMs can't be general intelligence or lack reasoning etc, to me seems like people witnessing the airplane and complaining that it isn't "real flying" because it isn't a bird flapping its wings (a large portion of the population held that point of view back then). It doesn't…

And look at the airplanes, they really can’t just land on a mountain slope or a tree without heavy maintenance afterwards. Those people weren’t all stupid, they questioned the promise of flying servicemen delivering mail or milk to their window and flying on a personal aircar to their workplace. Just like todays promises about whatever the CEOs telltales are. Imagining bullshit isn’t unique to this century. Aerospace…

This pretty much. Everyone knows that LLMs are great for text generation and processing. What people has been questioning is the end goals as promised by its builders, i.e. is it useful? And from most of what I saw, it's very much a toy.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#174

Efficiency is now key. ~=$3400 per single task to meet human performance on this benchmark is a lot. Also it shows the bullets as "ARC-AGI-TUNED", which makes me think they did some undisclosed amount of fine-tuning (eg. via the API they showed off last week), so even more compute went into this task. We can compare this roughly to a human doing ARC-AGI puzzles, where a human will take (high variance in my subjective…

> ~=$3400 per single task report says it is $17 per task, and $6k for whole dataset of 400 tasks.

That's the low-compute mode. In the plot at the top where they score 88%, O3 High (tuned) is ~3.4k

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#177

Efficiency is now key. ~=$3400 per single task to meet human performance on this benchmark is a lot. Also it shows the bullets as "ARC-AGI-TUNED", which makes me think they did some undisclosed amount of fine-tuning (eg. via the API they showed off last week), so even more compute went into this task. We can compare this roughly to a human doing ARC-AGI puzzles, where a human will take (high variance in my subjective…

> ~=$3400 per single task report says it is $17 per task, and $6k for whole dataset of 400 tasks.

That’s for the low-compute configuration that doesn’t reach human-level performance (not far though)

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#179

Let me go against some skeptics and explain why I think full o3 is pretty much AGI or at least embodies most essential aspects of AGI. What has been lacking so far in frontier LLMs is the ability to reliably deal with the right level of abstraction for a given problem. Reasoning is useful but often comes out lacking if one cannot reason at the right level of abstraction. (Note that many humans can't either when they…

Agree. AGI is here. I feel such a sense of pride in our species.

[deleted]

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#180
post #177

Earlier quoted context omitted.

> ~=$3400 per single task report says it is $17 per task, and $6k for whole dataset of 400 tasks.

That’s for the low-compute configuration that doesn’t reach human-level performance (not far though)

I referred on high compute mode. They have table with breakdown here: https://arcprize.org/blog/oai-o3-pub-breakthrough
Post reply on HN