OpenAI O3 breakthrough high score on ARC-AGI-PUB
61–70 of 1001 posts
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#62So now not only are the models closed, but so are their evals?! This is a "semi-private" eval. WTH is that supposed to mean? I'm sure the model is great but I refuse to take their word for it.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#63Human performance is 85% [1]. o3 high gets 87.5%. This means we have an algorithm to get to human level performance on this task. If you think this task is an eval of general reasoning ability, we have an algorithm for that now. There's a lot of work ahead to generalize o3 performance to all domains. I think this explains why many researchers feel AGI is within reach, now that we have an algorithm that works. Congrat…
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#64My initial impression: it's very impressive and very exciting. My skeptical impression: it's complete hubris to conflate ARC or any benchmark with truly general intelligence. I know my skepticism here is identical to moving goalposts. More and more I am shifting my personal understanding of general intelligence as a phenomenon we will only ever be able to identify with the benefit of substantial retrospect. As it is…
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#65Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#66So now not only are the models closed, but so are their evals?! This is a "semi-private" eval. WTH is that supposed to mean? I'm sure the model is great but I refuse to take their word for it.
Thats how i understand it
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#67Human performance is 85% [1]. o3 high gets 87.5%. This means we have an algorithm to get to human level performance on this task. If you think this task is an eval of general reasoning ability, we have an algorithm for that now. There's a lot of work ahead to generalize o3 performance to all domains. I think this explains why many researchers feel AGI is within reach, now that we have an algorithm that works. Congrat…
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#68Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#69How much longer can I get paid $150k to write code ?
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#70My initial impression: it's very impressive and very exciting. My skeptical impression: it's complete hubris to conflate ARC or any benchmark with truly general intelligence. I know my skepticism here is identical to moving goalposts. More and more I am shifting my personal understanding of general intelligence as a phenomenon we will only ever be able to identify with the benefit of substantial retrospect. As it is…
These comments are getting ridiculous. I remember when this test was first discussed here on HN and everyone agreed that it clearly proves current AI models are not "intelligent" (whatever that means). And people tried to talk me down when I theorised this test will get nuked soon - like all the ones before. It's time people woke up and realised that the old age of AI is over. This new kind is here to stay and it wil…