Live data from Hacker News

OpenAI o3 and o4-mini

openai.com

31–40 of 527 posts

Re: OpenAI o3 and o4-mini

#31
post #27

Earlier quoted context omitted.

[flagged]

Some people don't blindly trust the marketing department of the publisher

Then it doesn't even matter what they name the model since it's just marketing that they wouldn't trust anyway.

Re: OpenAI o3 and o4-mini

#32
Very impressive! But under arguably the most important benchmark -- SWE-bench verified for real-world coding tasks -- Claude 3.7 still remains the champion.[1]

Incredible how resilient Claude models have been for best-in-coding class.

[1] But by only about 1%, and inclusive of Claude's "custom scaffold" augmentation (which in practice I assume almost no one uses?). The new OpenAI models might still be effectively best in class now (or likely beating Claude with similar augmentation?).

Re: OpenAI o3 and o4-mini

#33
post #14

What is wrong with OpenAI? The naming of their models seems like it is intentionally confusing - maybe to distract from lack of progress? Honestly, I have no idea which model to use for simply everyday tasks anymore.

Seems to me like they're somewhat trying to simplify now.

GPT-N.m -> Non-reasoning

oN -> Reasoning

oN+1-mini -> Reasoning but speedy; cut-down version of an upcoming oN model (unclear if true or marketing)

It would be nice if they actually stick to this pattern.

Re: OpenAI o3 and o4-mini

#35

  ChatGPT Plus, Pro, and Team users will see o3, o4-mini, and o4-mini-high in the model selector starting today, replacing o1, o3‑mini, and o3‑mini‑high.
I subscribe to pro but don't yet see the new models (either in the Android app or on the web version).

Re: OpenAI o3 and o4-mini

#36

As a consumer, it is so exhausting keeping up with what model I should or can be using for the task I want to accomplish.

I think it can be confusing if you're just reading the news. If you use ChatGPT, the model selector has good brief explanations and teaching you about newly available options if you don't visit the dropdown. Anthropic does similarly.

Re: OpenAI o3 and o4-mini

#37
post #6

The pace of notable releases across the industry right now is unlike any time I remember since I started doing this in the early 2000's. And it feels like it's accelerating

Not really. We’re definitely in the incremental improvement stage at this point. Certainly no indication that progress is “accelerating”.

Re: OpenAI o3 and o4-mini

#38

As a consumer, it is so exhausting keeping up with what model I should or can be using for the task I want to accomplish.

[flagged]

"good at advanced reasoning", "fast at advanced reasoning", "slower at advanced reasoning but more advanced than the good one but not as fast but cant search the internet", "great at code and logic", "good for everyday tasks but awful at everything else", "faster for most questions but answers them incorrectly", "can draw but cant search", "can search but cant draw", "good for writing and doing creative things"

Re: OpenAI o3 and o4-mini

#39
Maybe OpenAI needs an easy mode for all these people saying 5 choices of models (and that's only if you pay) is simply too confusing for them.

They even provide a description in the UI of each before you select it, and it defaults to a model for you.

If you just want an answer of what you should use and can't be bothered to research them, just use o3(4)-mini and call it a day.

Post reply on HN