Earlier quoted context omitted.
I think OpenAI really needs to rethink its product naming, especially now that they have a portfolio where there's no such clear hierarchy, but they have a place along different axis (speed, cost, reasoning, capabilities, etc). Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.
Can't wait for the eventual rename to GPT Core, GPT Plus, GPT Pro, and GPT Pro Max models! I can see it now: > Unlock our industry leading reasoning features by upgrading to the GPT 4 Pro Max plan.
OpenAI O3-Mini
321–330 of 944 posts
Re: OpenAI O3-Mini
#322Earlier quoted context omitted.
That's such a counter-productive and frankly dumb thing to do. Just don't vote on them.
You have to pick one to continue the chat.
Re: OpenAI O3-Mini
#323Re: OpenAI O3-Mini
#324> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…
Re: OpenAI O3-Mini
#325Re: OpenAI O3-Mini
#326So far, it seems like this is the hierarchy o1 > GPT-4o > o3-mini > o1-mini > GPT-4o-mini o3 mini system card: https://cdn.openai.com/o3-mini-system-card.pdf
Re: OpenAI O3-Mini
#327Earlier quoted context omitted.
It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.
That would be 12%, why would you assume that is eaten by statistical noise?
Re: OpenAI O3-Mini
#328Re: OpenAI O3-Mini
#329Earlier quoted context omitted.
Yes I'd bet most users just 50/50 it, which actually makes it more remarkable that there was a 56% selection rate
I read the one on the left but choose the shorter one. The interface wastes so much screen real estate already and the answers are usually overly verbose unless I've given explicit instructions on how to answer.
Re: OpenAI O3-Mini
#330Earlier quoted context omitted.
Funny - I had ChatGPT document some stuff for me this week and asked which responses I preferred as well. Didn’t bother reading either of them, just selected one and went on with my day. If it were me I would have set up a “hey do you mind if we give you two results and you can pick your favorite?” prompt to weed out people like me.
I'm surprised how many people claim to do this. You can just not select one.