Earlier quoted context omitted.
Sonnet and Gemini saw fairly substantial perf increases recenly
Love Sonnet but 3.7 is not obviously an improvement over 3.5 in my real world usage. Gemini 2.5 pro is great, has replaced most others for me (Grok I use for things that require realtime answers)
OpenAI o3 and o4-mini
391–400 of 527 posts
Re: OpenAI o3 and o4-mini
#392Earlier quoted context omitted.
Not really. We’re definitely in the incremental improvement stage at this point. Certainly no indication that progress is “accelerating”.
ChatGPT 3 : iPhone 1 A bunch of models later, we're about on the iPhone 4-5 now. Feels about right.
Re: OpenAI o3 and o4-mini
#393Earlier quoted context omitted.
This misses the point. LLMs will do things like move a knight by a single square as if it were a pawn. Chess is an extremely well understood game, and the rules about how things move is almost certainly well-represented in the training data. These models cannot even make legal chess moves. That’s incredibly basic logic, and it shows how LLMs are still completely incapable of reasoning or understanding. Many kinds of…
>These models cannot even make legal chess moves. That’s incredibly basic logic, and it shows how LLMs are still completely incapable of reasoning or understanding. Yeah they can. There's a link I shared to prove it which you've conveniently ignored. LLMs learn by predicting, failing and getting a little better, rinse and repeat. Pre-training is not like reading a book. LLMs trained on chess games play chess just fin…
Re: OpenAI o3 and o4-mini
#394The pace of notable releases across the industry right now is unlike any time I remember since I started doing this in the early 2000's. And it feels like it's accelerating
Re: OpenAI o3 and o4-mini
#395As usual, it's a frustrating experience for anything more complex than the usual problems everyone else does.
Re: OpenAI o3 and o4-mini
#396Earlier quoted context omitted.
having many models from the same company in some haphazard strategy doesn't equate to "industry fragmentation". it's just confusion
OpenAI's continued growth and press coverage relative to their peers leads to me to believe it isn't *just* confusion, even if it is confusing.
Re: OpenAI o3 and o4-mini
#397As a consumer, it is so exhausting keeping up with what model I should or can be using for the task I want to accomplish.
Gemini 2.5 Pro for every single task was the meta until this release. Will have to reassess now.
Re: OpenAI o3 and o4-mini
#398Re: OpenAI o3 and o4-mini
#399Some interesting hallucinations going on here!
Re: OpenAI o3 and o4-mini
#400Ok, I’m a bit underwhelmed. I’ve asked it a fairly technical question, about a very niche topic (Final Fantasy VII reverse engineering): https://chatgpt.com/share/68001766-92c8-8004-908f-fb185b7549... With right knowledge and web searches one can answer this question in a matter of minutes at most. The model fumbled around modding forums and other sites and did manage to find some good information but then started to…