Live data from Hacker News

OpenAI o3 and o4-mini

openai.com

241–250 of 527 posts

Re: OpenAI o3 and o4-mini

#241
post #120

Earlier quoted context omitted.

Im old enough to remember the mystery and hype before o*/o1/strawberry that was supposed to be essentially AGI. We had serious news outlets write about senior people at OpenAI quitting because o1 was SkyNet Now we're up to o4, AGI is still not even in near site (depending on your definition, I know). And OpenAI is up to about 5000 employees. I'd think even before AGI a new model would be able to cover for at least 45…

Meanwhile even the highest ranked models can’t do simple logic tasks. GothamChess on YouTube did some tests where he played against a bunch of the best models and every single one of them failed spectacularly. They’d happily lose a queen to take a pawn. They failed to understand how pieces are even allowed to move, hallucinated the existence of new pieces, repeatedly declared checkmate when it wasn’t, etc. I tried it…

Chess is not exactly a simple logic task. It requires you to keep track of 32 things in a 2d space.

I remember being extremely surprised when I could ask GPT3 to rotate a 3d model of a car in it's head and ask it about what I would see when sitting inside, or which doors would refuse to open because they're in contact with the ground.

It really depends on how much you want to shift the goalposts on what constitutes "simple".

Re: OpenAI o3 and o4-mini

#242
post #4

Where's the comparison with Gemini 2.5 Pro?

For coding, I like the Aider polyglot benchmark, since it covers multiple programming languages. Gemini 2.5 Pro got 72.9% o3 high gets 81.3%, o4-mini high gets 68.9%

Isn't it easy to train on the specific Exercism exercises that this benchmark uses?

Re: OpenAI o3 and o4-mini

#243

Earlier quoted context omitted.

I find Gemini 2.5 truly remarkable and overall better than Claude, which I was a big fan of

Still doesn't work well in Cursor unfortunately.

Works well in RA.Aid --in fact I'd recommend it as the default model in terms of overall cost and capability.

Re: OpenAI o3 and o4-mini

#244

Earlier quoted context omitted.

Yes, however Claude advertised 70.3%[1] on SWE bench verified when using the following scaffolding: > For Claude 3.7 Sonnet and Claude 3.5 Sonnet (new), we use a much simpler approach with minimal scaffolding, where the model decides which commands to run and files to edit in a single session. Our main “no extended thinking” pass@1 result simply equips the model with the two tools described here—a bash tool, and a fi…

I think you may have misread the footnote. That simpler setup results in the 62.3%/63.7% score. The 70.3% score results from a high-compute parallel setup with rejection sampling and ranking: > For our “high compute” number we adopt additional complexity and parallel test-time compute as follows: > We sample multiple parallel attempts with the scaffold above > We discard patches that break the visible regression test…

Somehow completely missed that, thanks!

I think reading this makes it even clearer that the 70.3% score should just be discarded from the benchmarks. "I got a 7%-8% higher SWE benchmark score by doing a bunch of extra work and sampling a ton of answers" is not something a typical user is going to have already set up when logging onto Claude and asking it a SWE style question.

Personally, it seems like an illegitimate way to juice the numbers to me (though Claude was transparent with what they did so it's all good, and it's not uninteresting to know you can boost your score by 8% with the right tooling).

Re: OpenAI o3 and o4-mini

#245
post #80

To plan a visit to a dark sky place, I used duck.ai (Duckduckgo's experimental AI chat feature) to ask five different AIs on what date the new moon will happen in August 2025. GPT-4o mini: The new moon in August 2025 will occur on August 12. Llama 3.3 70B: The new moon in August 2025 is expected to occur on August 16, 2025. Claude 3 Haiku: The new moon in August 2025 will occur on August 23, 2025. o3-mini: Based on a…

"Who was the President of the United States when Neil Armstrong walked on the moon?" Gemini 2.5 refuses to answer this because it is too political.

I call bs on this: https://g.co/gemini/share/ed38e9d38b02

Re: OpenAI o3 and o4-mini

#246
post #87

So at this point OpenAI has 6 reasoning models, 4 flagship chat models, and 7 cost optimized models. So that's 17 models in total and that's not even counting their older models and more specialized ones. Compare this with Anthropic that has 7 models in total and 2 main ones that they promote. This is just getting to be a bit much, seems like they are trying to cover for the fact that they haven't actually done much.…

"haven't actually done much" being popularizing the chat llm and absolutely dwarfing the competition in paid usage

Relative to the hype they've been spinning to attract investment, casting the launch and commercialization of ChatGPT as their greatest achievement really is a quite significant downgrade, especially given that they really only got there first because they were the first entity reckless enough to deploy such a tool to the public.

It's easy to forget what smart, connected people were saying about how AI would evolve by ~a year ago, when in fact what we've gotten since then is a whole bunch of diminishing returns and increasingly sketchy benchmark shenanigans. I have no idea when a real AGI breakthrough will happen, but if you're a person who wants it to happen (I am not), you have to admit to yourself that the last year or so has been disappointing---even if you won't admit it to anybody else.

Re: OpenAI o3 and o4-mini

#247
post #148

The most annoying part of all this is they replaced o1 with o3 without any notices or warnings. This is why I hate proprietary models.

Meanwhile we have people elsewhere in the thread complaining about too many models. Assuming OpenAI are correct that o3 is strictly an improvement over o1 then I don't see why they'd keep o1 around. When they upgrade gpt-o4 they don't let you use the old version, after all.

>Assuming OpenAI are correct that o3 is strictly an improvement over o1 then I don't see why they'd keep o1 around.

Imagine if every time your favorite SaaS had an update, they renamed the product. Yesterday you were using Slack S7, and today you're suddenly using Slack 9S-o. That was fine in the desktop era, when new releases happened once a year - not every few weeks. You just can't keep up with all the versions.

I think they should just stick with one brand and announce new releases as just incremental updates to that same brand/product (even if the underlying models are different): "the DeepSearch Update" or "The April 2025 Reasoning Update" etc.

The model picker should be replaced entirely with a router that automatically detects which underlying model to use. Power users could have optional checkboxes like "Think harder" or "Code mode" as settings, if they want to guide the router toward more specialized models.

Re: OpenAI o3 and o4-mini

#248
post #87

Earlier quoted context omitted.

"haven't actually done much" being popularizing the chat llm and absolutely dwarfing the competition in paid usage

I guess it was related to the last period, rather than the full picture

What are people expecting here honestly? This thread is ridiculous.

Re: OpenAI o3 and o4-mini

#249
post #195

Earlier quoted context omitted.

They're rumored to be working on a social network to rival X with the focus being on image generations. https://techcrunch.com/2025/04/15/openai-is-reportedly-devel... The play now seems to be less AGI, more "too big to fail" / use all the capital to morph into a FAANG bigtech. My bet is that they'll develop a suite of office tools that leverage their model, chat/communication tools, a browser, and perhaps a device.…

I appreciate the info and I have a question: Why would anyone use a social network run by Sam Altman? No offense but his reputation is chaotic neutral to say the least. Social networks require a ton of momentum to get going. BlueSky already ate all the momentum that X lost.

Most people don't care about techies or tech drama. They just use the platforms their friends do.

ChatGPT images are the biggest thing on social media right now. My wife is turning photos of our dogs into people. There's a new GPT4o meme trending on TikTok every day. Using GPT4o as the basis of a social media network could be just the kickstart a new social media platform needs.

Re: OpenAI o3 and o4-mini

#250
post #195

Earlier quoted context omitted.

They're rumored to be working on a social network to rival X with the focus being on image generations. https://techcrunch.com/2025/04/15/openai-is-reportedly-devel... The play now seems to be less AGI, more "too big to fail" / use all the capital to morph into a FAANG bigtech. My bet is that they'll develop a suite of office tools that leverage their model, chat/communication tools, a browser, and perhaps a device.…

I appreciate the info and I have a question: Why would anyone use a social network run by Sam Altman? No offense but his reputation is chaotic neutral to say the least. Social networks require a ton of momentum to get going. BlueSky already ate all the momentum that X lost.

Social networks have to be the most chaotic neutral thing ever made. It's like, "hey everyone! Come share what ever you want on my servers!"
Post reply on HN