Live data from Hacker News

OpenAI o3 and o4-mini

openai.com

121–130 of 527 posts

Re: OpenAI o3 and o4-mini

#121

So at this point OpenAI has 6 reasoning models, 4 flagship chat models, and 7 cost optimized models. So that's 17 models in total and that's not even counting their older models and more specialized ones. Compare this with Anthropic that has 7 models in total and 2 main ones that they promote. This is just getting to be a bit much, seems like they are trying to cover for the fact that they haven't actually done much.…

To use that criticism for this release ain't really fair, as these will replace the old models (o3 will replace o1, o4-mini will replace o3-mini). On a more general level - sure, but they aren't planning to use this release to add a larger number of models, it's just that deprecating/killing the old models can't be done overnight.

As someone who doesn't use anything OpenAI (for all the reasons), I have to agree with the GP. It's all baffling. Why is there an o3-mini and an o4-mini? Why on earth are there so many models?

Once you get to this point you're putting the paradox of choice on the user - I used to use a particular brand toothpaste for years until it got to the point where I'd be in the supermarket looking at a wall of toothpaste all by the same brand with no discernible difference between the products. Why is one of them called "whitening"? Do the others not do that? Why is this one called "complete" and that one called "complete ultra"? That would suggest that the "complete" one wasn't actually complete. I stopped using that brand of toothpaste as it become impossible to know which was the right product within the brand.

If I was assessing the AI landscape today, where the leading models are largely indistinguishable in day to day use, I'd look at OpenAI's wall of toothpaste and immediately discount them.

Re: OpenAI o3 and o4-mini

#123

Earlier quoted context omitted.

Main advantage over Sonnet is Gemini 2.5 doesn't try to make a bunch of unrelated changes like it's rewriting my project from scratch.

Also that Gemini 2.5 still doesn’t support prompt caching, which is huge for tools like Cline.

2.5 Pro supports prompt caching now: https://cloud.google.com/vertex-ai/generative-ai/docs/models...

Re: OpenAI o3 and o4-mini

#125

A suggestion for OpenAI to create more meaningful model names: {Size}-{Quarter/Year}-{Speed/Accuracy}-{Specialty} Where: * Size is XS/S/M/L/XL/XXL to indicate overall capability level * Quarter/Year like Q2-25 * Speed/Accuracy indicated as Fast/Balanced/Precise * Optional specialty tag like Code/Vision/Science/etc Example model names: * L-Q2-25-Fast-Code (Large model from Q2 2025, optimized for speed, specializes in…

I think they should name them after fictional characters. Bonus points if they're trademarked characters.

"You gotta try Mickey, it beats the crap out of Gandalf in coding."

Re: OpenAI o3 and o4-mini

#126

Earlier quoted context omitted.

I don't know how anyone could look at any of this and say ponderously: it's basically the same as Nov 2022 ChatGPT. Thus strategically they're pivoting to social to become too big to fail.

I mean, it's not fucking AGI/ASI. No amount of LLM flip floppery is going to get us terminators. If this starts looking differently and the pace picks up, I won't be giving analysis on OpenAI anymore. I'll start packing for the hills. But to OpenAI's credit, I also don't see how minting another FAANG isn't an incredible achievement. Like - wow - this tech giant was willed into existence. Can't we marvel at that a lit…

I don't know what AGI/ASI means to you.

I'm bullish on the models, and my first quiet 5 minutes after the announcement was spent thinking how many of the people I walked past days would be different if the computer Just Did It(tm) (I don't think their day would be different, so I'm not bullish on ASI-even-if-achieved, I guess?)

I think binary analysis that flips between "this is a propped up failure, like when banks get bailouts" and "I'd run away from civilization" isn't really worth much.

Re: OpenAI o3 and o4-mini

#127

Very impressive! But under arguably the most important benchmark -- SWE-bench verified for real-world coding tasks -- Claude 3.7 still remains the champion.[1] Incredible how resilient Claude models have been for best-in-coding class. [1] But by only about 1%, and inclusive of Claude's "custom scaffold" augmentation (which in practice I assume almost no one uses?). The new OpenAI models might still be effectively bes…

Also, if you're using Cursor AI, it seems to have much better integration with Claude where it can reflect on its own things and go off and run commands. I don't see it doing that with Gemini or the O1 models.

Re: OpenAI o3 and o4-mini

#128
post #112
post #21

It's pretty frustrating to see a press release with "Try on ChatGPT" and then not see the models available even though I'm paying them $200/mo.

They are all now available on the Pro plan. Y'all really ought to have a little bit more grace to wait 30 minutes after the announcement for the rollout.

Or maybe OpenAI could wait until they'd released it before telling people to use it now.

Re: OpenAI o3 and o4-mini

#129
post #46

4o and o4 at the same time. Excellent work on the product naming, whoever did that.

It took me reading your comment to realize that they were different and this wasn’t deja vu. Maybe that says more about me than OpenAI, but my gut agrees with you.

Re: OpenAI o3 and o4-mini

#130
post #123

Earlier quoted context omitted.

Also that Gemini 2.5 still doesn’t support prompt caching, which is huge for tools like Cline.

2.5 Pro supports prompt caching now: https://cloud.google.com/vertex-ai/generative-ai/docs/models...

Oh, that must’ve been in the last few days. Weird that it’s only in 2.5 Pro preview but at least they’re headed in the right direction.

Now they just need a decent usage dashboard that doesn’t take a day to populate or require additional GCP monitoring services to break out the model usage.

Post reply on HN