Not sure what the goal is with Codex CLI. It's not running a local LLM right, just a CLI to make API calls from the terminal?
This might be their answer to claude code more than anything else.
OpenAI o3 and o4-mini
51–60 of 527 posts
Re: OpenAI o3 and o4-mini
#52Earlier quoted context omitted.
[flagged]
"good at advanced reasoning", "fast at advanced reasoning", "slower at advanced reasoning but more advanced than the good one but not as fast but cant search the internet", "great at code and logic", "good for everyday tasks but awful at everything else", "faster for most questions but answers them incorrectly", "can draw but cant search", "can search but cant draw", "good for writing and doing creative things"
Re: OpenAI o3 and o4-mini
#53Very impressive! But under arguably the most important benchmark -- SWE-bench verified for real-world coding tasks -- Claude 3.7 still remains the champion.[1] Incredible how resilient Claude models have been for best-in-coding class. [1] But by only about 1%, and inclusive of Claude's "custom scaffold" augmentation (which in practice I assume almost no one uses?). The new OpenAI models might still be effectively bes…
[0] swebench.com/#verified
Re: OpenAI o3 and o4-mini
#54What is wrong with OpenAI? The naming of their models seems like it is intentionally confusing - maybe to distract from lack of progress? Honestly, I have no idea which model to use for simply everyday tasks anymore.
Re: OpenAI o3 and o4-mini
#55Re: OpenAI o3 and o4-mini
#56What is wrong with OpenAI? The naming of their models seems like it is intentionally confusing - maybe to distract from lack of progress? Honestly, I have no idea which model to use for simply everyday tasks anymore.
Re: OpenAI o3 and o4-mini
#57What is wrong with OpenAI? The naming of their models seems like it is intentionally confusing - maybe to distract from lack of progress? Honestly, I have no idea which model to use for simply everyday tasks anymore.
Seems to me like they're somewhat trying to simplify now. GPT-N.m -> Non-reasoning oN -> Reasoning oN+1-mini -> Reasoning but speedy; cut-down version of an upcoming oN model (unclear if true or marketing) It would be nice if they actually stick to this pattern.
Re: OpenAI o3 and o4-mini
#58This is just getting to be a bit much, seems like they are trying to cover for the fact that they haven't actually done much. All these models feel like they took the exact same base model, tweaked a few things and released it as an entirely new model rather than updating the existing ones. In fact based on some of the other comments here it sounds like these are just updates to their existing model, but they release them as new models to create more media buzz.
Re: OpenAI o3 and o4-mini
#59So at this point OpenAI has 6 reasoning models, 4 flagship chat models, and 7 cost optimized models. So that's 17 models in total and that's not even counting their older models and more specialized ones. Compare this with Anthropic that has 7 models in total and 2 main ones that they promote. This is just getting to be a bit much, seems like they are trying to cover for the fact that they haven't actually done much.…
Re: OpenAI o3 and o4-mini
#60OpenAI be like: o1, o1-mini, o1-pro, o3, o4-mini, gpt-4, gpt-4o, gpt-4-turbo, gpt-4.5, gpt-4.1, gpt-4o-mini, gpt-4.1-mini, gpt-4.1-nano, gpt-3.5-turbo