I’ll take the China Deluxe instead, actually. I’ve been incredibly pleased with DeepSeek this past week. Wonderful product, I love seeing its brain when it’s thinking.
OpenAI O3-Mini
71–80 of 944 posts
Re: OpenAI O3-Mini
#72> While OpenAI o1 remains our broader general knowledge reasoning model, OpenAI o3-mini provides a specialized alternative for technical domains requiring precision and speed. I feel like this naming scheme is growing a little tired. o1 is for general knowledge reasoning, o3-mini replaces o1-mini but might be more specialized than o1 for certain technical domains...the "o" in "4o" is for "omni" (referring to its mult…
They really need someone in marketing. If the model is for technical stuff, then call it the technical model. How is anyone supposed to know what these model names mean? The only page of theirs attempting to explain this is a total disaster. https://platform.openai.com/docs/models
...I like "DALL·E" and "Whisper" as names a lot, though, FWIW :p
Re: OpenAI O3-Mini
#73why should anyone use this when deepseek is free/cheaper? openai is no longer relevant.
I don't think OpenAI is training on your data. At least they say they don't, and I believe that. I wouldn't be surprised if the NSA or something has access to data if they request it or something though. But DeepSeek clearly states in their terms of service that they can train on your API data or use it for other purposes. Which one might assume their government can access as well. We need direct eval comparisons bet…
Like they said they were committed to being “open”?
Re: OpenAI O3-Mini
#74Anyone else confused by inconsistency in performance numbers between this announcement and the concurrent system card? https://cdn.openai.com/o3-mini-system-card.pdf For example- GPQA diamond system card: o1-preview 0.68 GPQA diamond PR release: o1-preview 0.78 Also, how should we interpret the 3 different shading colors in the barplots (white, dotted, heavy dotted on top of white)...
Re: OpenAI O3-Mini
#75> While OpenAI o1 remains our broader general knowledge reasoning model, OpenAI o3-mini provides a specialized alternative for technical domains requiring precision and speed. I feel like this naming scheme is growing a little tired. o1 is for general knowledge reasoning, o3-mini replaces o1-mini but might be more specialized than o1 for certain technical domains...the "o" in "4o" is for "omni" (referring to its mult…
Re: OpenAI O3-Mini
#76Did anyone else notice that o3-mini's SWE bench dropped from 61% in the leaked System Card earlier today to 49.3% in this blog post, which puts o3-mini back in line with Claude on real-world coding tasks? Am I missing something?
Maybe they found a need to quantize it further for release, or lobotomise it with more "alignment".
Re: OpenAI O3-Mini
#77Earlier quoted context omitted.
I think OpenAI really needs to rethink its product naming, especially now that they have a portfolio where there's no such clear hierarchy, but they have a place along different axis (speed, cost, reasoning, capabilities, etc). Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.
It's like AWS SKU naming (`c5d.metal`, `p5.48xlarge`, etc.), except non-technical consumers are expected to understand it.
Re: OpenAI O3-Mini
#78Earlier quoted context omitted.
They really need someone in marketing. If the model is for technical stuff, then call it the technical model. How is anyone supposed to know what these model names mean? The only page of theirs attempting to explain this is a total disaster. https://platform.openai.com/docs/models
> They really need someone in marketing. Who said this is not intentional? It seems to work well given that people are hyped every time there's a release, no matter how big the actual improvements are — I'm pretty sure "o3-mini" works better for that purpose than "GPT 4.1.3"
Why would the marketing team of all people call it GPT 4.1.3?
Re: OpenAI O3-Mini
#79Re: OpenAI O3-Mini
#80Anyone else confused by inconsistency in performance numbers between this announcement and the concurrent system card? https://cdn.openai.com/o3-mini-system-card.pdf For example- GPQA diamond system card: o1-preview 0.68 GPQA diamond PR release: o1-preview 0.78 Also, how should we interpret the 3 different shading colors in the barplots (white, dotted, heavy dotted on top of white)...