The pace of notable releases across the industry right now is unlike any time I remember since I started doing this in the early 2000's. And it feels like it's accelerating
Not really. We’re definitely in the incremental improvement stage at this point. Certainly no indication that progress is “accelerating”.
OpenAI o3 and o4-mini
111–120 of 527 posts
Re: OpenAI o3 and o4-mini
#112It's pretty frustrating to see a press release with "Try on ChatGPT" and then not see the models available even though I'm paying them $200/mo.
Re: OpenAI o3 and o4-mini
#113I'm starting to be reminded of the razor blade business.
Re: OpenAI o3 and o4-mini
#114Earlier quoted context omitted.
Gemini 2.5 Pro is widely considered superior to 3.7 Sonnet now by heavy users, but they don't have an SWE-bench score. Shows that looking at one such benchmark isn't very telling. Main advantage over Sonnet being that it's better at using a large amount of context, which is enormously helpful during coding tasks. Sonnet is still an incredibly impressive model as it held the crown for 6 months, which may as well be a…
Main advantage over Sonnet is Gemini 2.5 doesn't try to make a bunch of unrelated changes like it's rewriting my project from scratch.
Re: OpenAI o3 and o4-mini
#115When are they going to release o3-high? I don't think it's in the API, and I certainly don't see it in the web app (Pro).
Re: OpenAI o3 and o4-mini
#116Very impressive! But under arguably the most important benchmark -- SWE-bench verified for real-world coding tasks -- Claude 3.7 still remains the champion.[1] Incredible how resilient Claude models have been for best-in-coding class. [1] But by only about 1%, and inclusive of Claude's "custom scaffold" augmentation (which in practice I assume almost no one uses?). The new OpenAI models might still be effectively bes…
I haven't been following them that closely, but are people finding these benchmarks relevant? It seems like these companies could just tune their models to do well on particular benchmarks
Re: OpenAI o3 and o4-mini
#117Earlier quoted context omitted.
ChatGPT was released in 2022, so OP's point stands perfectly well.
They're rumored to be working on a social network to rival X with the focus being on image generations. https://techcrunch.com/2025/04/15/openai-is-reportedly-devel... The play now seems to be less AGI, more "too big to fail" / use all the capital to morph into a FAANG bigtech. My bet is that they'll develop a suite of office tools that leverage their model, chat/communication tools, a browser, and perhaps a device.…
Re: OpenAI o3 and o4-mini
#118To the extent that reasoning is noisy and models can go astray during it, this helps inject truth back into the reasoning loop.
Is there some well known equivalent to Moores Law for token use? We're headed in a direction where LLM control loops can run 24/7 generating tokens to reason about live sensor data, and calling tools to act on it.
Re: OpenAI o3 and o4-mini
#119Earlier quoted context omitted.
They're rumored to be working on a social network to rival X with the focus being on image generations. https://techcrunch.com/2025/04/15/openai-is-reportedly-devel... The play now seems to be less AGI, more "too big to fail" / use all the capital to morph into a FAANG bigtech. My bet is that they'll develop a suite of office tools that leverage their model, chat/communication tools, a browser, and perhaps a device.…
I don't know how anyone could look at any of this and say ponderously: it's basically the same as Nov 2022 ChatGPT. Thus strategically they're pivoting to social to become too big to fail.
If this starts looking differently and the pace picks up, I won't be giving analysis on OpenAI anymore. I'll start packing for the hills.
But to OpenAI's credit, I also don't see how minting another FAANG isn't an incredible achievement. Like - wow - this tech giant was willed into existence. Can't we marvel at that a little bit without worrying about LLMs doing our taxes?
Re: OpenAI o3 and o4-mini
#120So at this point OpenAI has 6 reasoning models, 4 flagship chat models, and 7 cost optimized models. So that's 17 models in total and that's not even counting their older models and more specialized ones. Compare this with Anthropic that has 7 models in total and 2 main ones that they promote. This is just getting to be a bit much, seems like they are trying to cover for the fact that they haven't actually done much.…
Now we're up to o4, AGI is still not even in near site (depending on your definition, I know). And OpenAI is up to about 5000 employees. I'd think even before AGI a new model would be able to cover for at least 4500 of those employees being fired, is that not the case?