Live data from Hacker News

OpenAI O3-Mini

openai.com

321–330 of 944 posts

Re: OpenAI O3-Mini

#321
post #12

Earlier quoted context omitted.

I think OpenAI really needs to rethink its product naming, especially now that they have a portfolio where there's no such clear hierarchy, but they have a place along different axis (speed, cost, reasoning, capabilities, etc). Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.

Can't wait for the eventual rename to GPT Core, GPT Plus, GPT Pro, and GPT Pro Max models! I can see it now: > Unlock our industry leading reasoning features by upgrading to the GPT 4 Pro Max plan.

Oh, I'll probably wait for GPT 4 Pro Max v2 NG (improved)

Re: OpenAI O3-Mini

#322
post #295

Earlier quoted context omitted.

That's such a counter-productive and frankly dumb thing to do. Just don't vote on them.

You have to pick one to continue the chat.

Why not always pick the one on the left, for example? I understand wanting to speed through and not spend time doing labor for OpenAI, but it seems counter-productive to spend any time feeding it false information.

Re: OpenAI O3-Mini

#323
post #174
post #134

Earlier quoted context omitted.

Even the full model scores below Claude on livebench so a distilled version will likely be even worse.

Based on the leaderboard R1 is significantly better than Claude? https://livebench.ai/#/

Not at coding.

Re: OpenAI O3-Mini

#324

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

[deleted]

Re: OpenAI O3-Mini

#325
O3-mini solved this prompt. DeepSeek R1 had a mental breakdown. The prompt: “Bob is facing forward. To his left is Ann, to his right is Cathy. Ann and Cathy are facing backwards. Who is on Ann’s left?”

Re: OpenAI O3-Mini

#326
post #3

So far, it seems like this is the hierarchy o1 > GPT-4o > o3-mini > o1-mini > GPT-4o-mini o3 mini system card: https://cdn.openai.com/o3-mini-system-card.pdf

no the reasoning models should not directly be compared with the normal models: they often take 10 times as long to answer which only makes sense for difficult questions

Re: OpenAI O3-Mini

#327

Earlier quoted context omitted.

It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.

That would be 12%, why would you assume that is eaten by statistical noise?

The OPs comment is probably a testament of that. With such a poorly designed A/B test I doubt this has a p-value of < 0.10.

Re: OpenAI O3-Mini

#329

Earlier quoted context omitted.

Yes I'd bet most users just 50/50 it, which actually makes it more remarkable that there was a 56% selection rate

I read the one on the left but choose the shorter one. The interface wastes so much screen real estate already and the answers are usually overly verbose unless I've given explicit instructions on how to answer.

The default level of verbosity you get without explicitly prompting for it to be succinct makes me think there’s an office full of workers getting paid by the token.

Re: OpenAI O3-Mini

#330
post #307

Earlier quoted context omitted.

Funny - I had ChatGPT document some stuff for me this week and asked which responses I preferred as well. Didn’t bother reading either of them, just selected one and went on with my day. If it were me I would have set up a “hey do you mind if we give you two results and you can pick your favorite?” prompt to weed out people like me.

I'm surprised how many people claim to do this. You can just not select one.

I think it’s somewhat natural and am not personally surprised. It’s easy to quickly select an option, that has no consequence, compared to actively considering that not selecting something is an option. Not selecting something feels more like actively participating than just checking a box and moving on. /shrug
Post reply on HN