Live data from Hacker News

OpenAI O3 breakthrough high score on ARC-AGI-PUB

arcprize.org

21–30 of 1001 posts

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#22

Congratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far. A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning. It's obvious to everyone that…

Are there any single-step non-reasoner models that do well on this benchmark?

I wonder how well the latest Claude 3.5 Sonnet does on this benchmark and if it's near o1.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#23
post #8

How much longer can I get paid $150k to write code ?

Frontier expert specialist programmers will always be in demand.

Generalist junior and senior engineers will need to think of a different career path in less than 5 years as more layoffs will reduce the software engineering workforce.

It looks like it may be the way things are if progress in the o1, o3, oN models and other LLMs continues on.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#24
O3 High (tuned) model scored an 88% at what looks like $6,000/task haha

I think soon we'll be pricing any kind of tasks by their compute costs. So basically, human = $50/task, AI = $6,000/task, use human. If AI beats human, use AI? Ofc that's considering both get 100% scores on the task

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#26

This is also wildly ahead in SWE-bench (71.7%, previous 48%) and Frontier Math (25% on high compute, previous 2%). So much for a plateau lol.

I legit see that if there is not even a new breakthrough just one week, people start shouting plateau plateau.. Our rate of progress is extraordinary and any downplay of it seems like stupid

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#27
post #8

How much longer can I get paid $150k to write code ?

Often what happens is the golf-course phenomenon. As golfing gets less popular, low and mid tier golf courses go out of business as they simply aren't needed. But at the same time demand for high end golf courses actually skyrockets because people who want to golf either can give it up or go higher end. This I think will happen with programmers. Rote programming will slowly die out, while demand for super high end wi…

Where does this golf-course phenomenon come from? It doesn't really match the real world or how golfing works.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#28

This is also wildly ahead in SWE-bench (71.7%, previous 48%) and Frontier Math (25% on high compute, previous 2%). So much for a plateau lol.

>Frontier Math (25% on high compute, previous 2%)

This is so insane that I can't help but be skeptical. I know FM answer key is private, but they have to send the questions to OpenAI in order to score the models. And a significant jump on this benchmark sure would increase a company's valuation...

Happy to be wrong on this.

Post reply on HN