Live data from Hacker News

OpenAI O3-Mini

openai.com

261–270 of 944 posts

Re: OpenAI O3-Mini

#261
post #12

Earlier quoted context omitted.

I think OpenAI really needs to rethink its product naming, especially now that they have a portfolio where there's no such clear hierarchy, but they have a place along different axis (speed, cost, reasoning, capabilities, etc). Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.

They can still do models o3o, oo3 and 3oo. Mini-o3o-high, not to be confused with mini-O3o-high (the first o is capital).

You’re thinking too small. What about o10, O1o, o3-m1n1?

Re: OpenAI O3-Mini

#262
post #28
post #12

Earlier quoted context omitted.

I think OpenAI really needs to rethink its product naming, especially now that they have a portfolio where there's no such clear hierarchy, but they have a place along different axis (speed, cost, reasoning, capabilities, etc). Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.

It's like AWS SKU naming (`c5d.metal`, `p5.48xlarge`, etc.), except non-technical consumers are expected to understand it.

Those are not names but hashes used to look up the specs.

Re: OpenAI O3-Mini

#263

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

Yes I'd bet most users just 50/50 it, which actually makes it more remarkable that there was a 56% selection rate

I read the one on the left but choose the shorter one.

The interface wastes so much screen real estate already and the answers are usually overly verbose unless I've given explicit instructions on how to answer.

Re: OpenAI O3-Mini

#264

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

[deleted]

Re: OpenAI O3-Mini

#265

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

they also pay contractors to do these evaluations with much more detailed metrics, no idea which their number is based on though

Re: OpenAI O3-Mini

#267
post #20

Earlier quoted context omitted.

There's no moat, and they have to work even harder. Competition is good.

I really don't think this is true. OpenAI has no moat because they have nothing unique; they're using mostly other people's (like Transformers) architectures and other companies hardware. Their value-prop (moat) is that they've burnt more money than everybody else. That moat is trivially circumvented by lighting a larger pile of money and less trivially by lighting the pile more efficently. OpenAI isn't the only comp…

> That moat is trivially circumvented by lighting a larger pile of money and less trivially by lighting the pile more efficently.

Google with all its money and smart engineers was not able to build a simple chat application.

Re: OpenAI O3-Mini

#268
Does anyone know the current usage limits for o3-mini and o3-mini-high when used through the ChatGPT interface? I tried to find them on the OpenAI Knowledgebase, but couldn’t find anything about that.

Re: OpenAI O3-Mini

#270

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something.

Because when you think about it, it really is quite damning. Minus statistical noise it's no better.

Post reply on HN