So, next step in reasoning is open world reasoning now?
OpenAI O3 breakthrough high score on ARC-AGI-PUB
21–30 of 1001 posts
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#22Congratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far. A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning. It's obvious to everyone that…
I wonder how well the latest Claude 3.5 Sonnet does on this benchmark and if it's near o1.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#23How much longer can I get paid $150k to write code ?
Generalist junior and senior engineers will need to think of a different career path in less than 5 years as more layoffs will reduce the software engineering workforce.
It looks like it may be the way things are if progress in the o1, o3, oN models and other LLMs continues on.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#24I think soon we'll be pricing any kind of tasks by their compute costs. So basically, human = $50/task, AI = $6,000/task, use human. If AI beats human, use AI? Ofc that's considering both get 100% scores on the task
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#25Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#26This is also wildly ahead in SWE-bench (71.7%, previous 48%) and Frontier Math (25% on high compute, previous 2%). So much for a plateau lol.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#27How much longer can I get paid $150k to write code ?
Often what happens is the golf-course phenomenon. As golfing gets less popular, low and mid tier golf courses go out of business as they simply aren't needed. But at the same time demand for high end golf courses actually skyrockets because people who want to golf either can give it up or go higher end. This I think will happen with programmers. Rote programming will slowly die out, while demand for super high end wi…
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#28This is also wildly ahead in SWE-bench (71.7%, previous 48%) and Frontier Math (25% on high compute, previous 2%). So much for a plateau lol.
This is so insane that I can't help but be skeptical. I know FM answer key is private, but they have to send the questions to OpenAI in order to score the models. And a significant jump on this benchmark sure would increase a company's valuation...
Happy to be wrong on this.
Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB
#29Great. Now we have to think of a new way to move the goalposts.