Live data from Hacker News

OpenAI O3-Mini

openai.com

451–460 of 944 posts

Re: OpenAI O3-Mini

#452

Earlier quoted context omitted.

That's a great idea. In your frontend, do you write in the same text entry field as the bot? I use oobabooga/text-generation-webui and I findit's a little awkward to edit the bot responses.

No, but the chat divs are all contenteditable.

Oh! That is an excellent solution. I wish it was that easy in every UI.

Re: OpenAI O3-Mini

#453
post #235

Earlier quoted context omitted.

It's nice to see Google finally having competition in a space it used to really dominate (though they definitely still are holding their own with all the Gemini naming). I feel like it takes real effort to have product names be this confusing and capricious

Gemini naming seems pretty straightforward at this point. 2.0 is the full model, flash is a smaller/faster/cheaper model, and flash thinking is a smaller/faster/cheaper reasoning model with Cost.

> 2.0 is the full model

Not quite. "2.0 Flash" is also called 2.0. The "Pro" models are the full models. But, I love how they have both "gemini-exp-1206" and "gemini-2.0-flash-thinking-exp-01-21". The first one doesn't even say what type of model it is, presumably it should have been "gemini-2.0-pro-exp-1206", but they didn't want to label it that for some reason, and now they're putting a hyphen in the date string where they weren't before.

Not to mention they have both "Flash" and "Flash-8B"... which I think will confuse people. IMO, it should be "Flash-${Parameters}B" for both of them if they're going to mention it for one.

But, I generally think Google's Gemini naming structure has been pretty decent.

Re: OpenAI O3-Mini

#454

Earlier quoted context omitted.

The OPs comment is probably a testament of that. With such a poorly designed A/B test I doubt this has a p-value of < 0.10.

Erm, why not? A 0.56 result with n=1000 ratings is statistically significantly better than 0.5 with a p-value of 0.00001864, well beyond any standard statistical significance threshold I've ever heard of. I don't know how many ratings they collected but 1000 doesn't seem crazy at all. Assuming of course that raters are blind to which model is which and the order of the 2 responses is randomized with every rating -- o…

  > If so, where do they indicate they failed to randomize/blind the raters?

  Win rate if user is under time constraint
This is hard to read tbh. Is it STEM? Non-STEM? If it is STEM then this shows there is a bias. If it is Non-STEM then this shows a bias. If it is a mix, well we can't know anything without understanding the split.

Note that Non-STEM is still within error. STEM is less than 2 sigma variance, so our confidence still shouldn't be that high.

Re: OpenAI O3-Mini

#456

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

You know you can configure default instructions to your prompts, right?

I have something like “always be terse and blunt with your answers.”

Re: OpenAI O3-Mini

#457
post #412

Earlier quoted context omitted.

It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.

And another way to rephrase it is that almost half of the users prefer the older model, which is terrible PR.

Not if the goal is to claim that the models deliver comparable quality, but with the new one excelling at something else (here: inferrence cost).

Re: OpenAI O3-Mini

#458

O3-mini solved this prompt. DeepSeek R1 had a mental breakdown. The prompt: “Bob is facing forward. To his left is Ann, to his right is Cathy. Ann and Cathy are facing backwards. Who is on Ann’s left?”

Really interesting puzzle. I tried the different Chatgpts and they all solved it. I tried Deepseek and Deepseek R1, and Deepseek did have this meltdown:

Let's break down the problem step by step to understand the relationships and positions of Bob, Ann, and Cathy. 1. Understanding the Initial Setup

    Bob is facing forward.

    This means Bob's front is oriented in a particular direction, which we'll consider as the reference point for "forward."

    To his left is Ann, to his right is Cathy.

    If Bob is facing forward, then:

        Ann is positioned to Bob's left.

        Cathy is positioned to Bob's right.

    Ann and Cathy are facing backwards.

    Both Ann and Cathy are oriented in the opposite direction to Bob. If Bob is facing forward, then Ann and Cathy are facing backward.
2. Visualizing the Positions

To better understand the scenario, let's visualize the positions: Copy

Forward Direction: ↑

Bob (facing forward) | | Ann (facing backward) | / | / | / | / | / | / | / |/ |

And then only the character | in a newline forever.

Re: OpenAI O3-Mini

#459
post #340

Earlier quoted context omitted.

It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.

Yeah. I immediately thought: I wonder if that 56% is in one or two categories and the rest are worse?

44% of the people prefers the existing model ?

Re: OpenAI O3-Mini

#460
post #71

I’ll take the China Deluxe instead, actually. I’ve been incredibly pleased with DeepSeek this past week. Wonderful product, I love seeing its brain when it’s thinking.

I recently tried Gemini-1.5-Pro for the first time. It was clearly better than DeepSeek or any of the OpenAI models available to Plus subscribers.

Try https://deepmind.google/technologies/gemini/flash-thinking/
Post reply on HN