So at this point OpenAI has 6 reasoning models, 4 flagship chat models, and 7 cost optimized models. So that's 17 models in total and that's not even counting their older models and more specialized ones. Compare this with Anthropic that has 7 models in total and 2 main ones that they promote. This is just getting to be a bit much, seems like they are trying to cover for the fact that they haven't actually done much.…
"haven't actually done much" being popularizing the chat llm and absolutely dwarfing the competition in paid usage
OpenAI o3 and o4-mini
101–110 of 527 posts
Re: OpenAI o3 and o4-mini
#102Surprisingly, they didn't provide a comparison to Sonnet 3.7 or Gemini Pro 2.5—probably because, while both are impressive, they're only slightly better by comparison. Lets see what the pricing looks like.
Re: OpenAI o3 and o4-mini
#103Re: OpenAI o3 and o4-mini
#104Earlier quoted context omitted.
"haven't actually done much" being popularizing the chat llm and absolutely dwarfing the competition in paid usage
ChatGPT was released in 2022, so OP's point stands perfectly well.
https://techcrunch.com/2025/04/15/openai-is-reportedly-devel...
The play now seems to be less AGI, more "too big to fail" / use all the capital to morph into a FAANG bigtech.
My bet is that they'll develop a suite of office tools that leverage their model, chat/communication tools, a browser, and perhaps a device.
They're going to try to turn into Google (with maybe a bit of Apple and Meta) before Google turns into them.
Near-term, I don't see late stage investors as recouping their investment. But in time, this may work out well for them. There's a tremendous amount of inefficiency and lack of competition amongst the big tech players. They've been so large that nobody else could effectively challenge them. Now there's a "startup" with enough capital to start eating into big tech's more profitable business lines.
Re: OpenAI o3 and o4-mini
#105Earlier quoted context omitted.
Gemini 2.5 Pro is widely considered superior to 3.7 Sonnet now by heavy users, but they don't have an SWE-bench score. Shows that looking at one such benchmark isn't very telling. Main advantage over Sonnet being that it's better at using a large amount of context, which is enormously helpful during coding tasks. Sonnet is still an incredibly impressive model as it held the crown for 6 months, which may as well be a…
Main advantage over Sonnet is Gemini 2.5 doesn't try to make a bunch of unrelated changes like it's rewriting my project from scratch.
Re: OpenAI o3 and o4-mini
#106So at this point OpenAI has 6 reasoning models, 4 flagship chat models, and 7 cost optimized models. So that's 17 models in total and that's not even counting their older models and more specialized ones. Compare this with Anthropic that has 7 models in total and 2 main ones that they promote. This is just getting to be a bit much, seems like they are trying to cover for the fact that they haven't actually done much.…
Re: OpenAI o3 and o4-mini
#107So at this point OpenAI has 6 reasoning models, 4 flagship chat models, and 7 cost optimized models. So that's 17 models in total and that's not even counting their older models and more specialized ones. Compare this with Anthropic that has 7 models in total and 2 main ones that they promote. This is just getting to be a bit much, seems like they are trying to cover for the fact that they haven't actually done much.…
The old Chinese strategy of having 7343 different phone models with almost the same specs to confuse the customer better
Re: OpenAI o3 and o4-mini
#108This reminds me of keeping up with all the latest JavaScript framework trivia circa the ~2010s
https://krausest.github.io/js-framework-benchmark/2025/table...
Re: OpenAI o3 and o4-mini
#109So at this point OpenAI has 6 reasoning models, 4 flagship chat models, and 7 cost optimized models. So that's 17 models in total and that's not even counting their older models and more specialized ones. Compare this with Anthropic that has 7 models in total and 2 main ones that they promote. This is just getting to be a bit much, seems like they are trying to cover for the fact that they haven't actually done much.…
"haven't actually done much" being popularizing the chat llm and absolutely dwarfing the competition in paid usage
Re: OpenAI o3 and o4-mini
#110Very impressive! But under arguably the most important benchmark -- SWE-bench verified for real-world coding tasks -- Claude 3.7 still remains the champion.[1] Incredible how resilient Claude models have been for best-in-coding class. [1] But by only about 1%, and inclusive of Claude's "custom scaffold" augmentation (which in practice I assume almost no one uses?). The new OpenAI models might still be effectively bes…
Claude got 63.2% according to the swebench.com leaderboard (listed as "Tools + Claude 3.7 Sonnet (2025-02-24)).[0] OpenAI said they got 69.1% in their blog post. [0] swebench.com/#verified