Live data from Hacker News

OpenAI O3-Mini

openai.com

481–490 of 944 posts

Re: OpenAI O3-Mini

#481
post #412

Earlier quoted context omitted.

It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.

And another way to rephrase it is that almost half of the users prefer the older model, which is terrible PR.

Typically in these tests you have three options "A is better", "B is better" or "they're equal/can't decide". So if 56% prefer O3 Mini, it's likely that way less than half prefer O1.also, the way I understand it, they're comparing a mini model with a large one.

Re: OpenAI O3-Mini

#483

Earlier quoted context omitted.

That seems very bad. What's the point of a new model that's worse than 4o? I guess it's cheaper in the API and a bit better at coding - but, this doesn't seem compelling. With DeepSeek I heard OpenAI saying the plan was to move releases on models that were meaningfully better than the competition. Seems like what we're getting is the scheduled releases that are worse than the current versions.

It's quite a bit better than coding --- they hint that it can tie o1's performance for coding, which already benchmarks higher than 4o. And it's significantly cheaper, and presumably faster. I believe API costs account for the vast majority of COGS at most today's AI startups, so they would be very motivated to switch to a cheaper model that has similar performance.

Right. For large-volume requests that use reasoning this will be quite useful. I have a task that requires the LLM to convert thousands of free-text statements into SQL select statements, and o3-mini-high is able to get many of the more complicated ones that GPT-4o and Sonnet 3.5 failed at. So I will be switching this task to either o3-mini or DeepSeek-R1.

Re: OpenAI O3-Mini

#484

I’ll take the China Deluxe instead, actually. I’ve been incredibly pleased with DeepSeek this past week. Wonderful product, I love seeing its brain when it’s thinking.

Being able to see the thinking trace in R1 is so useful, as you can go back and see if it's getting stuck, making a wrong assumption, missing data, etc. To me that makes it materially more useful than the OpenAI reasoning models, which seem impressive, but are much harder to inspect/debug.

It's almost like watching a stoned centipede having a panic attack about moving its legs. It also makes it obvious that these models (not just R1 I suppose) need to learn some kind of priority estimation to stop overthinking irrelevant issues and leave them to the normal token prediction, while focusing on the stuff that matters.

Nevertheless, R1's reasoning chains are already shorter in tokens than o1's while having similar results, and apparently o3-mini's too.

Re: OpenAI O3-Mini

#485
post #333

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

This prompt is like "See Attendant" on the gas pump. I'm just going to use another AI instead for this chat.

Glad to know I’m not the only person who just drives to the next station when I see a “see attendant” message.

Re: OpenAI O3-Mini

#486
post #412

Earlier quoted context omitted.

And another way to rephrase it is that almost half of the users prefer the older model, which is terrible PR.

Typically in these tests you have three options "A is better", "B is better" or "they're equal/can't decide". So if 56% prefer O3 Mini, it's likely that way less than half prefer O1.also, the way I understand it, they're comparing a mini model with a large one.

If you use ChatGPT, it sometimes gives you two versions of its response, and you have to choose one or the other if you want to continue prompting. Sure, not picking a response might be a third category. But if that's how they were approaching the analysis, they could have put out a more favorable-looking stat.

Re: OpenAI O3-Mini

#487
post #43

Haven't used openai in a bit -- whyyy did they change "system" role (now basically an industry-wide standard) to "developer"? That seems pointlessly disruptive.

2 years ago I'd say it's an oversight, because there's 0 chance a top down directive would ask for this.

But given how OpenAI employees act online these days I wouldn't be surprised if someone on the ground proposed it as a way to screw with all the 3rd parties who are using OpenAI compatible endpoints or even use OpenAI's SDK in their official docs in some cases.

Re: OpenAI O3-Mini

#488

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

Also, it's not clear if the preference comes from the quality of the 'meat' of the answer, or the way it reports its thinking and the speed with which it responds. With o1, I get a marked feeling of impatience waiting for it to spit something out, and the 'progress of thought' is in faint grey text I can't read. With o3, the 'progress of thought' comes quickly, with more to read, and is more engaging even if I don't actually get anything more than entertainment value.

I'm not going to say there's nothing substantive about o3 vs. o1, but I absolutely do not put it past Sam Altman to juice the stats every chance he gets.

Re: OpenAI O3-Mini

#489

Does anyone know the current usage limits for o3-mini and o3-mini-high when used through the ChatGPT interface? I tried to find them on the OpenAI Knowledgebase, but couldn’t find anything about that.

For Plus users the limits are:

o3-mini-high: 50 messages per week (just like o1, but it seems like these are non-shared limits, so you can have 50 messages per week with o1, run out, and still have 50 messages with o3-mini-high to use)

o3-mini: 150 messages per day

Source for the latter is their press release. They were more vague about o3-mini-high, but people have already tested its limits just by using it, and got the pop-up for 25 messages left after sending 25 messages.

It's nice not to worry about running out of o1 messages now and have a faster model that's mostly as good (potentially better in some areas?). OpenAI really needs to release a middle tier for 30 to $40 though that has the same models as Pro but without infinite usage. I hate not having the smartest model and I don't want to pay $200; there's probably a middle ground where they can make as much or more money from me on a subscription tier that gives limited access to o1-pro.

Re: OpenAI O3-Mini

#490

Does anyone know the current usage limits for o3-mini and o3-mini-high when used through the ChatGPT interface? I tried to find them on the OpenAI Knowledgebase, but couldn’t find anything about that.

[deleted]
Post reply on HN