Live data from Hacker News

OpenAI O3-Mini

openai.com

71–80 of 944 posts

Re: OpenAI O3-Mini

#71

I’ll take the China Deluxe instead, actually. I’ve been incredibly pleased with DeepSeek this past week. Wonderful product, I love seeing its brain when it’s thinking.

I recently tried Gemini-1.5-Pro for the first time. It was clearly better than DeepSeek or any of the OpenAI models available to Plus subscribers.

Re: OpenAI O3-Mini

#72

> While OpenAI o1 remains our broader general knowledge reasoning model, OpenAI o3-mini provides a specialized alternative for technical domains requiring precision and speed. I feel like this naming scheme is growing a little tired. o1 is for general knowledge reasoning, o3-mini replaces o1-mini but might be more specialized than o1 for certain technical domains...the "o" in "4o" is for "omni" (referring to its mult…

They really need someone in marketing. If the model is for technical stuff, then call it the technical model. How is anyone supposed to know what these model names mean? The only page of theirs attempting to explain this is a total disaster. https://platform.openai.com/docs/models

Ugh, and some of the rows of that table are "sets of models" while some are singular models...there's the "Flagship models" section at the top only for "GPT models" to be heralded as "Our fast, versatile, high intelligence flagship models" in the NEXT section...

...I like "DALL·E" and "Whisper" as names a lot, though, FWIW :p

Re: OpenAI O3-Mini

#73
post #35
post #23

why should anyone use this when deepseek is free/cheaper? openai is no longer relevant.

I don't think OpenAI is training on your data. At least they say they don't, and I believe that. I wouldn't be surprised if the NSA or something has access to data if they request it or something though. But DeepSeek clearly states in their terms of service that they can train on your API data or use it for other purposes. Which one might assume their government can access as well. We need direct eval comparisons bet…

> I don't think OpenAI is training on your data. At least they say they don't, and I believe that.

Like they said they were committed to being “open”?

Re: OpenAI O3-Mini

#74

Anyone else confused by inconsistency in performance numbers between this announcement and the concurrent system card? https://cdn.openai.com/o3-mini-system-card.pdf For example- GPQA diamond system card: o1-preview 0.68 GPQA diamond PR release: o1-preview 0.78 Also, how should we interpret the 3 different shading colors in the barplots (white, dotted, heavy dotted on top of white)...

[deleted]

Re: OpenAI O3-Mini

#75

> While OpenAI o1 remains our broader general knowledge reasoning model, OpenAI o3-mini provides a specialized alternative for technical domains requiring precision and speed. I feel like this naming scheme is growing a little tired. o1 is for general knowledge reasoning, o3-mini replaces o1-mini but might be more specialized than o1 for certain technical domains...the "o" in "4o" is for "omni" (referring to its mult…

Inscrutable naming is a proven strategy for muddying the waters.

Re: OpenAI O3-Mini

#76

Did anyone else notice that o3-mini's SWE bench dropped from 61% in the leaked System Card earlier today to 49.3% in this blog post, which puts o3-mini back in line with Claude on real-world coding tasks? Am I missing something?

Maybe they found a need to quantize it further for release, or lobotomise it with more "alignment".

Or the number was never real to begin with.

Re: OpenAI O3-Mini

#77
post #28
post #12

Earlier quoted context omitted.

I think OpenAI really needs to rethink its product naming, especially now that they have a portfolio where there's no such clear hierarchy, but they have a place along different axis (speed, cost, reasoning, capabilities, etc). Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.

It's like AWS SKU naming (`c5d.metal`, `p5.48xlarge`, etc.), except non-technical consumers are expected to understand it.

Have you seen Azure VM SKU naming? It's.. impressive.

Re: OpenAI O3-Mini

#78
post #64

Earlier quoted context omitted.

They really need someone in marketing. If the model is for technical stuff, then call it the technical model. How is anyone supposed to know what these model names mean? The only page of theirs attempting to explain this is a total disaster. https://platform.openai.com/docs/models

> They really need someone in marketing. Who said this is not intentional? It seems to work well given that people are hyped every time there's a release, no matter how big the actual improvements are — I'm pretty sure "o3-mini" works better for that purpose than "GPT 4.1.3"

> I'm pretty sure "o3-mini" works better for that purpose than "GPT 4.1.3"

Why would the marketing team of all people call it GPT 4.1.3?

Re: OpenAI O3-Mini

#80

Anyone else confused by inconsistency in performance numbers between this announcement and the concurrent system card? https://cdn.openai.com/o3-mini-system-card.pdf For example- GPQA diamond system card: o1-preview 0.68 GPQA diamond PR release: o1-preview 0.78 Also, how should we interpret the 3 different shading colors in the barplots (white, dotted, heavy dotted on top of white)...

Actually sounds like benchslop to me.
Post reply on HN