OpenAI o3 and o4-mini
91–100 of 527 posts
Re: OpenAI o3 and o4-mini
#92Very impressive! But under arguably the most important benchmark -- SWE-bench verified for real-world coding tasks -- Claude 3.7 still remains the champion.[1] Incredible how resilient Claude models have been for best-in-coding class. [1] But by only about 1%, and inclusive of Claude's "custom scaffold" augmentation (which in practice I assume almost no one uses?). The new OpenAI models might still be effectively bes…
Claude got 63.2% according to the swebench.com leaderboard (listed as "Tools + Claude 3.7 Sonnet (2025-02-24)).[0] OpenAI said they got 69.1% in their blog post. [0] swebench.com/#verified
Re: OpenAI o3 and o4-mini
#93Very impressive! But under arguably the most important benchmark -- SWE-bench verified for real-world coding tasks -- Claude 3.7 still remains the champion.[1] Incredible how resilient Claude models have been for best-in-coding class. [1] But by only about 1%, and inclusive of Claude's "custom scaffold" augmentation (which in practice I assume almost no one uses?). The new OpenAI models might still be effectively bes…
Re: OpenAI o3 and o4-mini
#94Very impressive! But under arguably the most important benchmark -- SWE-bench verified for real-world coding tasks -- Claude 3.7 still remains the champion.[1] Incredible how resilient Claude models have been for best-in-coding class. [1] But by only about 1%, and inclusive of Claude's "custom scaffold" augmentation (which in practice I assume almost no one uses?). The new OpenAI models might still be effectively bes…
Claude got 63.2% according to the swebench.com leaderboard (listed as "Tools + Claude 3.7 Sonnet (2025-02-24)).[0] OpenAI said they got 69.1% in their blog post. [0] swebench.com/#verified
> For Claude 3.7 Sonnet and Claude 3.5 Sonnet (new), we use a much simpler approach with minimal scaffolding, where the model decides which commands to run and files to edit in a single session. Our main “no extended thinking” pass@1 result simply equips the model with the two tools described here—a bash tool, and a file editing tool that operates via string replacements—as well as the “planning tool” mentioned above in our TAU-bench results.
Arguably this shouldn't be counted though?
[1] https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-...
Re: OpenAI o3 and o4-mini
#95I assume this announcement is all 256k, while the base model 4.1 just shot up this week to a million.
Re: OpenAI o3 and o4-mini
#96Re: OpenAI o3 and o4-mini
#97As a consumer, it is so exhausting keeping up with what model I should or can be using for the task I want to accomplish.
Re: OpenAI o3 and o4-mini
#98Earlier quoted context omitted.
Gemini 2.5 Pro is widely considered superior to 3.7 Sonnet now by heavy users, but they don't have an SWE-bench score. Shows that looking at one such benchmark isn't very telling. Main advantage over Sonnet being that it's better at using a large amount of context, which is enormously helpful during coding tasks. Sonnet is still an incredibly impressive model as it held the crown for 6 months, which may as well be a…
Main advantage over Sonnet is Gemini 2.5 doesn't try to make a bunch of unrelated changes like it's rewriting my project from scratch.
Re: OpenAI o3 and o4-mini
#99A suggestion for OpenAI to create more meaningful model names: {Size}-{Quarter/Year}-{Speed/Accuracy}-{Specialty} Where: * Size is XS/S/M/L/XL/XXL to indicate overall capability level * Quarter/Year like Q2-25 * Speed/Accuracy indicated as Fast/Balanced/Precise * Optional specialty tag like Code/Vision/Science/etc Example model names: * L-Q2-25-Fast-Code (Large model from Q2 2025, optimized for speed, specializes in…
Re: OpenAI o3 and o4-mini
#100So at this point OpenAI has 6 reasoning models, 4 flagship chat models, and 7 cost optimized models. So that's 17 models in total and that's not even counting their older models and more specialized ones. Compare this with Anthropic that has 7 models in total and 2 main ones that they promote. This is just getting to be a bit much, seems like they are trying to cover for the fact that they haven't actually done much.…
"haven't actually done much" being popularizing the chat llm and absolutely dwarfing the competition in paid usage