Earlier quoted context omitted.
It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.
And another way to rephrase it is that almost half of the users prefer the older model, which is terrible PR.
OpenAI O3-Mini
481–490 of 944 posts
Re: OpenAI O3-Mini
#482Re: OpenAI O3-Mini
#483Earlier quoted context omitted.
That seems very bad. What's the point of a new model that's worse than 4o? I guess it's cheaper in the API and a bit better at coding - but, this doesn't seem compelling. With DeepSeek I heard OpenAI saying the plan was to move releases on models that were meaningfully better than the competition. Seems like what we're getting is the scheduled releases that are worse than the current versions.
It's quite a bit better than coding --- they hint that it can tie o1's performance for coding, which already benchmarks higher than 4o. And it's significantly cheaper, and presumably faster. I believe API costs account for the vast majority of COGS at most today's AI startups, so they would be very motivated to switch to a cheaper model that has similar performance.
Re: OpenAI O3-Mini
#484I’ll take the China Deluxe instead, actually. I’ve been incredibly pleased with DeepSeek this past week. Wonderful product, I love seeing its brain when it’s thinking.
Being able to see the thinking trace in R1 is so useful, as you can go back and see if it's getting stuck, making a wrong assumption, missing data, etc. To me that makes it materially more useful than the OpenAI reasoning models, which seem impressive, but are much harder to inspect/debug.
Nevertheless, R1's reasoning chains are already shorter in tokens than o1's while having similar results, and apparently o3-mini's too.
Re: OpenAI O3-Mini
#485> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…
This prompt is like "See Attendant" on the gas pump. I'm just going to use another AI instead for this chat.
Re: OpenAI O3-Mini
#486Earlier quoted context omitted.
And another way to rephrase it is that almost half of the users prefer the older model, which is terrible PR.
Typically in these tests you have three options "A is better", "B is better" or "they're equal/can't decide". So if 56% prefer O3 Mini, it's likely that way less than half prefer O1.also, the way I understand it, they're comparing a mini model with a large one.
Re: OpenAI O3-Mini
#487Haven't used openai in a bit -- whyyy did they change "system" role (now basically an industry-wide standard) to "developer"? That seems pointlessly disruptive.
But given how OpenAI employees act online these days I wouldn't be surprised if someone on the ground proposed it as a way to screw with all the 3rd parties who are using OpenAI compatible endpoints or even use OpenAI's SDK in their official docs in some cases.
Re: OpenAI O3-Mini
#488> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…
I'm not going to say there's nothing substantive about o3 vs. o1, but I absolutely do not put it past Sam Altman to juice the stats every chance he gets.
Re: OpenAI O3-Mini
#489Does anyone know the current usage limits for o3-mini and o3-mini-high when used through the ChatGPT interface? I tried to find them on the OpenAI Knowledgebase, but couldn’t find anything about that.
o3-mini-high: 50 messages per week (just like o1, but it seems like these are non-shared limits, so you can have 50 messages per week with o1, run out, and still have 50 messages with o3-mini-high to use)
o3-mini: 150 messages per day
Source for the latter is their press release. They were more vague about o3-mini-high, but people have already tested its limits just by using it, and got the pop-up for 25 messages left after sending 25 messages.
It's nice not to worry about running out of o1 messages now and have a faster model that's mostly as good (potentially better in some areas?). OpenAI really needs to release a middle tier for 30 to $40 though that has the same models as Pro but without infinite usage. I hate not having the smartest model and I don't want to pay $200; there's probably a middle ground where they can make as much or more money from me on a subscription tier that gives limited access to o1-pro.
Re: OpenAI O3-Mini
#490Does anyone know the current usage limits for o3-mini and o3-mini-high when used through the ChatGPT interface? I tried to find them on the OpenAI Knowledgebase, but couldn’t find anything about that.