Live data from Hacker News

OpenAI O3 breakthrough high score on ARC-AGI-PUB

arcprize.org

991–1000 of 1001 posts

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#991

This is also wildly ahead in SWE-bench (71.7%, previous 48%) and Frontier Math (25% on high compute, previous 2%). So much for a plateau lol.

>Frontier Math (25% on high compute, previous 2%) This is so insane that I can't help but be skeptical. I know FM answer key is private, but they have to send the questions to OpenAI in order to score the models. And a significant jump on this benchmark sure would increase a company's valuation... Happy to be wrong on this.

viewed from a skeptical lens of incentives:

openai and epochai are both startups with every incentive to peddle this narrative. when no one else can independently verify.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#992
Am I understanding correctly, and the only thing with a bit of actual data released so far is the ARC-AGI piece from Francois Chollet? And every other claim has no further data released on it?

Serious question. I've browsed around, looked for the official release, but it seems to be just hear-say for now, except for the few little bits in the ARC-AGI article.

So some of the reactions seems quite far-fetched. I was quite amazed at first seeing the benchmarks, but then actually read the ARC-AGI article and a few other things about how it worked, learned a bit more about the different benchmarks, and realised we've no proper idea yet how o3 is working under the hood, the thing isn't even realeased.

It could be doing the same thing that chess-engines do except in several specific domains. Which would be very cool, but not necessarily "intelligent" or "generally intelligent" in any sense whatsoever! Will that kind of model lead to finding novel mathematical proofs, or actually "reasoning" or "thinking" in any way similar to a human, remains entirely uncertain.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#993
post #933

Earlier quoted context omitted.

I am not expert in llm reasoning but I think because of RL. You cannot use AlphaZero to play other games.

Nope. AlphaZero taught itself to play games like chess, shogi, and Go through self-play, starting from random moves. It was not given any strategies or human gameplay data but was provided with the basic rules of each game to guide its learning process.

Yes its reinforcement learning, but need to create policy and each policy is specialized for specific tasks.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#994
post #929

Earlier quoted context omitted.

I am not expert in llm reasoning but I think because of RL. You cannot use AlphaZero to play other games.

I thought that AlphaZero could play three games? Go, Chess and Shogi?

Think I mean Catan :)

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#995
For many people and businesses, navigating the frequently dangerous landscape of financial loss can be an intimidating and overwhelming process. Nevertheless, the knowledgeable staff at Wizard Hilton Cyber Tech provides a ray of hope and direction with their indispensable range of services. Their offerings are based on a profound grasp of the far-reaching and terrible effects that financial setbacks, whether they be the result of cyberattacks, data breaches, or other unforeseen tragedies, can have. Their highly-trained analysts work tirelessly to assess the scope of the damage, identifying the root causes and developing tailored strategies to mitigate the fallout. From recovering lost or corrupted data to restoring compromised systems and securing networks, Wizard Hilton Cyber Tech employs the latest cutting-edge technologies and industry best practices to help clients regain their financial footing. But their support goes beyond the technical realm, as their compassionate case managers provide a empathetic ear and practical advice to navigate the emotional and logistical challenges that often accompany financial upheaval. With a steadfast commitment to client success, Wizard Hilton Cyber Tech is a trusted partner in weathering the storm of financial loss, offering the essential services and peace of mind needed to emerge stronger and more resilient than before.

Re: OpenAI O3 breakthrough high score on ARC-AGI-PUB

#996

Earlier quoted context omitted.

I took a look at those examples that o3 can't solve. Looks similar to an IQ-test. Took me less time to figure out the 3 examples that it took to read your post. I was honestly a bit surprised to see how visual the tasks were. I had thought they were text based. So now I'm quite impressed that o3 can solve this type of task at all.

You must be a stem grad! Or perhaps an ensemble of Kaggle submissions?

I'm actually a panel of 10 randomly selected people.
Post reply on HN