The real heated contest here amongst the top AI labs is to see who can come up with the most confusing product names.
OpenAI O3-Mini
221–230 of 944 posts
Re: OpenAI O3-Mini
#222I think OpenAI should just have a single public facing "model" - all these names and versions are confusing. Imagine if Google, during it's accent, had a huge array of search engines with code names and notes about what it's doing behind the scenes. No, you open the page and type in box. If they can make it work better next month, great. (I understand this could not apply to developers or enterprise-type API usage).
Yelp suffered greatly in the early 2010s when Google started putting Google Maps listings (and their accompanying reviews) in their search results.
OpenAI will eventually unify their products as well.
Re: OpenAI O3-Mini
#223Earlier quoted context omitted.
How would the DeepSeek fit into this? Or can it not compare? I don't know much about this stuff, but I've heard recently many people talk about DeepSeek and how unexpected it was.
Deepseek V3 is equivalent to 4o. Deepseek R1 is equivalent to o1 (if not better) I think someone should just build an AI model comparing website at this point. Include all benchmarks and pricing
Re: OpenAI O3-Mini
#224I think OpenAI should just have a single public facing "model" - all these names and versions are confusing. Imagine if Google, during it's accent, had a huge array of search engines with code names and notes about what it's doing behind the scenes. No, you open the page and type in box. If they can make it work better next month, great. (I understand this could not apply to developers or enterprise-type API usage).
Thats the role of ChatGPT?
Re: OpenAI O3-Mini
#225Earlier quoted context omitted.
I think OpenAI really needs to rethink its product naming, especially now that they have a portfolio where there's no such clear hierarchy, but they have a place along different axis (speed, cost, reasoning, capabilities, etc). Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.
They're strongly tied to Microsoft, so confusing branding is to be expected.
- Xbox, Xbox 360, Xbox One, Xbox One S/X, Xbox Series S/X
- Windows 3.1...98, 2000, ME, XP, Vista, 7, 8, 10
I guess it's better than headphones names (QC35, WH-1000XM3, M50x, HD560s).
Re: OpenAI O3-Mini
#226Did anyone else notice that o3-mini's SWE bench dropped from 61% in the leaked System Card earlier today to 49.3% in this blog post, which puts o3-mini back in line with Claude on real-world coding tasks? Am I missing something?
The caption on the graph explains. > including with the open-source Agentless scaffold (39%) and an internal tools scaffold (61%), see our system card . I have no idea what an "internal tools scaffold" is but the graph on the card that they link directly to specifies "o3-mini (tools)" where the blog post is talking about others.
Instead of just generating a patch (copilot style), it generates the patch, applies the patch, runs the code, and then iterates based on the execution output.
Re: OpenAI O3-Mini
#227Earlier quoted context omitted.
> How is anyone supposed to know what these model names mean? Normies don't have to know - ChatGPT app focuses UX around capabilities and automatically picks the appropriate model for capabilities requested; you can see which model you're using and change it, but don't need to . As for the techies and self-proclaimed "AI experts" - OpenAI is the leader in the field, and one of the most well-known and talked about tec…
What if you use ASCII 234? Ω (edit: works!)
Re: OpenAI O3-Mini
#228Hopefully this is a big improvement from o1. o1 has been very disappointing after spending sufficient time with Claude Sonnet 3.5. It's like it actively tries to gaslight me and thinks it knows more than I do. It's too stubborn and confidently goes off in tangents, suggesting big changes to parts of the code that aren't the issue. Claude tends to be way better at putting the pieces together in its not-quite-mental-mo…
That said I often run into a sort of opposite issue with Claude. It's very good at making me feel like a genius. Sometimes I'll suggest trying a specific strategy or trying to define a concept on my own, and Claude enthusiastically agrees and takes us down a 2-3 hour rabbit hole that ends up being quite a waste of time for me to back track out of.
I'll then run a post-mortem through chatGPT and very often it points out the issue in my thinking very quickly.
That said I keep coming back to sonnet-3.5 for reasons I can't perfectly articulate. Perhaps because I like how it fluffs my ego lol. ChatGPT on the other hand feels a bit more brash. I do wonder if I should be using o1 as my daily driver.
I also don't have enough experience with o1 to determine if it would also take me down dead ends as well.
Re: OpenAI O3-Mini
#229Re: OpenAI O3-Mini
#230> While OpenAI o1 remains our broader general knowledge reasoning model, OpenAI o3-mini provides a specialized alternative for technical domains requiring precision and speed. I feel like this naming scheme is growing a little tired. o1 is for general knowledge reasoning, o3-mini replaces o1-mini but might be more specialized than o1 for certain technical domains...the "o" in "4o" is for "omni" (referring to its mult…
It's almost as bad as the Xbox naming scheme.