Live data from Hacker News

OpenAI o3 and o4-mini

openai.com

391–400 of 527 posts

Re: OpenAI o3 and o4-mini

#391

Earlier quoted context omitted.

Sonnet and Gemini saw fairly substantial perf increases recenly

Love Sonnet but 3.7 is not obviously an improvement over 3.5 in my real world usage. Gemini 2.5 pro is great, has replaced most others for me (Grok I use for things that require realtime answers)

Are you comparing it with or without thinking? I'd say it's a fairly big improvement in long thinking mode.

Re: OpenAI o3 and o4-mini

#392

Earlier quoted context omitted.

Not really. We’re definitely in the incremental improvement stage at this point. Certainly no indication that progress is “accelerating”.

ChatGPT 3 : iPhone 1 A bunch of models later, we're about on the iPhone 4-5 now. Feels about right.

It's more like GPT-3 is the Manchester Baby, and we're somewhere around IBM 700 series right now. Still a long way to go to iPhone, as much as the industry likes to pretend otherwise.

Re: OpenAI o3 and o4-mini

#393

Earlier quoted context omitted.

This misses the point. LLMs will do things like move a knight by a single square as if it were a pawn. Chess is an extremely well understood game, and the rules about how things move is almost certainly well-represented in the training data. These models cannot even make legal chess moves. That’s incredibly basic logic, and it shows how LLMs are still completely incapable of reasoning or understanding. Many kinds of…

>These models cannot even make legal chess moves. That’s incredibly basic logic, and it shows how LLMs are still completely incapable of reasoning or understanding. Yeah they can. There's a link I shared to prove it which you've conveniently ignored. LLMs learn by predicting, failing and getting a little better, rinse and repeat. Pre-training is not like reading a book. LLMs trained on chess games play chess just fin…

I think the point here is that if you have to pretrain it for every specific task, it's not artificial general intelligence, by definition.

Re: OpenAI o3 and o4-mini

#394
post #6

The pace of notable releases across the industry right now is unlike any time I remember since I started doing this in the early 2000's. And it feels like it's accelerating

How is this a notable release? It's strictly worse than Gemini 2.5 on coding &c, and only an iterative improvement over their own models. The only thing that struck me as particularly interesting was the native visual reasoning.

Re: OpenAI o3 and o4-mini

#395
I have been using o4-mini-high today. Most of the time for a file longer than 100 lines it stops generating randomly and won't complete a file unless I re-prompt it with the end of the missing file.

As usual, it's a frustrating experience for anything more complex than the usual problems everyone else does.

Re: OpenAI o3 and o4-mini

#396

Earlier quoted context omitted.

having many models from the same company in some haphazard strategy doesn't equate to "industry fragmentation". it's just confusion

OpenAI's continued growth and press coverage relative to their peers leads to me to believe it isn't *just* confusion, even if it is confusing.

I'd attribute that more to first mover advantage than a benefit from poor naming choices, though I do think they are likely to misattribute that to a causal relationship so that they keep doing the latter

Re: OpenAI o3 and o4-mini

#397

As a consumer, it is so exhausting keeping up with what model I should or can be using for the task I want to accomplish.

Gemini 2.5 Pro for every single task was the meta until this release. Will have to reassess now.

Mad tangent, but as an old timey MtG player it’s always jarring when someone uses “the meta” not to refer to the particular dynamics of their competitive ecosystem but to a single strategy within it. Impoverishes the concept, I feel, even in this case where I don’t actually think a single model is best at everything.

Re: OpenAI o3 and o4-mini

#398
o4 is doing a better job than o3 on my current project, and while this isn’t really a priority, its personality is somehow far more engaging now.

Re: OpenAI o3 and o4-mini

#400
post #298

Ok, I’m a bit underwhelmed. I’ve asked it a fairly technical question, about a very niche topic (Final Fantasy VII reverse engineering): https://chatgpt.com/share/68001766-92c8-8004-908f-fb185b7549... With right knowledge and web searches one can answer this question in a matter of minutes at most. The model fumbled around modding forums and other sites and did manage to find some good information but then started to…

Have you asked this same question to various other models out there in the wild? I am just curious if you have found some that performed better. I would ask some models myself, but I do not know the proper answer, so I would probably be gullible enough to believe whatever the various answers have in common.
Post reply on HN