Live data from Hacker News

OpenAI O3-Mini

openai.com

331–340 of 944 posts

Re: OpenAI O3-Mini

#332
post #20

Earlier quoted context omitted.

There's no moat, and they have to work even harder. Competition is good.

I really don't think this is true. OpenAI has no moat because they have nothing unique; they're using mostly other people's (like Transformers) architectures and other companies hardware. Their value-prop (moat) is that they've burnt more money than everybody else. That moat is trivially circumvented by lighting a larger pile of money and less trivially by lighting the pile more efficently. OpenAI isn't the only comp…

"OpenAI has no moat because they have nothing unique"

It seems they have high quality trainingsdata. And the knowledge to work with it.

Re: OpenAI O3-Mini

#333

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

This prompt is like "See Attendant" on the gas pump. I'm just going to use another AI instead for this chat.

Re: OpenAI O3-Mini

#334

O3-mini solved this prompt. DeepSeek R1 had a mental breakdown. The prompt: “Bob is facing forward. To his left is Ann, to his right is Cathy. Ann and Cathy are facing backwards. Who is on Ann’s left?”

R1 or R1-Distill? They are not the same thing. I think DeepSeek made a mistake releasing them at the same time and calling them all R1.

Full R1 solves this prompt easily for me.

Re: OpenAI O3-Mini

#335
o1-preview, o1, o1-mini, o3-mini, o3-mini (low), o3-mini (medium), o3-mini (high)...

What's next?

o4-mini (wet socks), o5-Eeny-meeny-miny-moe?

I thought they had a product manager over there.

They only need 2 names, right? ChatGPT and o.

ChatGPT-5 and o4 would be next.

This multiplication of the LLM loaves and fishes is kind of silly.

Re: OpenAI O3-Mini

#336
post #237

Earlier quoted context omitted.

Maybe they found a need to quantize it further for release, or lobotomise it with more "alignment".

> lobotomise Anyone can write very fast software if you don't mind it sometimes crashing or having weird bugs. Why do people try to meme as if AI is different? It has unexpected outputs sometimes, getting it to not do that is 50% "more alignment" and 50% "hallucinate less". Just today I saw someone get the Amazon bot to roleplay furry erotica. Funny, sure, but it's still obviously a bug that a *sales bot* would do th…

If somebody wants their Amazon bot to role play as an erotic furry, that’s up to them, right? Who cares. It is working as intended if it keeps them going back to the site and buying things I guess.

I don’t know why somebody would want that, seems annoying. But I also don’t expect people to explain why they do this kind of stuff.

Re: OpenAI O3-Mini

#337

Earlier quoted context omitted.

Have you considered the possibility that your feedback is used to choose what type of response to give to you specifically in the future? I would not consider purposely giving inaccurate feedback for this reason alone.

I don't want a model that's customized to my preferences. My preferences and understanding changes all the time. I want a single source model that's grounded in base truth. I'll let the model know how to structure it in my prompt.

Constang meh and fixing prompts to the right direction vs unable to escape the bubble

Re: OpenAI O3-Mini

#338
post #9

Earlier quoted context omitted.

What about "o1 Pro mode". Is that just o1 but with more reasoning time, like this new o3-mini's different amount of reasoning options?

o1-pro is a different model than o1.

Are you sure? Do you have any source for that? In this article[0] that was discussed here on HN this week, they say (claim):

> In fact, the O1 model used in OpenAI's ChatGPT Plus subscription for $20/month is basically the same model as the one used in the O1-Pro model featured in their new ChatGPT Pro subscription for 10x the price ($200/month, which raised plenty of eyebrows in the developer community); the main difference is that O1-Pro thinks for a lot longer before responding, generating vastly more COT logic tokens, and consuming a far larger amount of inference compute for every response.

Granted "basically" is pulling a lot of weight there, but that was the first time I'd seen anyone speculate either way.

[0] https://youtubetranscriptoptimizer.com/blog/05_the_short_cas...

Re: OpenAI O3-Mini

#339
Oh, sweet: both o3-mini low and high support integrated web search. No integrated web search with o1.

I prefer, for philosophical reasons, open weight and open process/science models, but OpenAI has done a very good job at productizing ChatGPT. I also use their 4o-mini API because it is cheap and compares well to using open models on Groq Cloud. I really love running local models with Ollama but the API venders keep the price so low that I understand most people not wanting the hasssle if running Deepseek-R, etc., locally.

Re: OpenAI O3-Mini

#340

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.

Yeah. I immediately thought: I wonder if that 56% is in one or two categories and the rest are worse?
Post reply on HN