Earlier quoted context omitted.
I took a quick look at the data and FWIW the votes look legit to me, if that's what you were wondering.
I'm fairly certain it was sarcasm.
OpenAI O3-Mini
451–460 of 944 posts
Re: OpenAI O3-Mini
#452Earlier quoted context omitted.
That's a great idea. In your frontend, do you write in the same text entry field as the bot? I use oobabooga/text-generation-webui and I findit's a little awkward to edit the bot responses.
No, but the chat divs are all contenteditable.
Re: OpenAI O3-Mini
#453Earlier quoted context omitted.
It's nice to see Google finally having competition in a space it used to really dominate (though they definitely still are holding their own with all the Gemini naming). I feel like it takes real effort to have product names be this confusing and capricious
Gemini naming seems pretty straightforward at this point. 2.0 is the full model, flash is a smaller/faster/cheaper model, and flash thinking is a smaller/faster/cheaper reasoning model with Cost.
Not quite. "2.0 Flash" is also called 2.0. The "Pro" models are the full models. But, I love how they have both "gemini-exp-1206" and "gemini-2.0-flash-thinking-exp-01-21". The first one doesn't even say what type of model it is, presumably it should have been "gemini-2.0-pro-exp-1206", but they didn't want to label it that for some reason, and now they're putting a hyphen in the date string where they weren't before.
Not to mention they have both "Flash" and "Flash-8B"... which I think will confuse people. IMO, it should be "Flash-${Parameters}B" for both of them if they're going to mention it for one.
But, I generally think Google's Gemini naming structure has been pretty decent.
Re: OpenAI O3-Mini
#454Earlier quoted context omitted.
The OPs comment is probably a testament of that. With such a poorly designed A/B test I doubt this has a p-value of < 0.10.
Erm, why not? A 0.56 result with n=1000 ratings is statistically significantly better than 0.5 with a p-value of 0.00001864, well beyond any standard statistical significance threshold I've ever heard of. I don't know how many ratings they collected but 1000 doesn't seem crazy at all. Assuming of course that raters are blind to which model is which and the order of the 2 responses is randomized with every rating -- o…
> If so, where do they indicate they failed to randomize/blind the raters?
Win rate if user is under time constraint
This is hard to read tbh. Is it STEM? Non-STEM? If it is STEM then this shows there is a bias. If it is Non-STEM then this shows a bias. If it is a mix, well we can't know anything without understanding the split.Note that Non-STEM is still within error. STEM is less than 2 sigma variance, so our confidence still shouldn't be that high.
Re: OpenAI O3-Mini
#455Re: OpenAI O3-Mini
#456> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…
I have something like “always be terse and blunt with your answers.”
Re: OpenAI O3-Mini
#457Earlier quoted context omitted.
It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.
And another way to rephrase it is that almost half of the users prefer the older model, which is terrible PR.
Re: OpenAI O3-Mini
#458O3-mini solved this prompt. DeepSeek R1 had a mental breakdown. The prompt: “Bob is facing forward. To his left is Ann, to his right is Cathy. Ann and Cathy are facing backwards. Who is on Ann’s left?”
Let's break down the problem step by step to understand the relationships and positions of Bob, Ann, and Cathy. 1. Understanding the Initial Setup
Bob is facing forward.
This means Bob's front is oriented in a particular direction, which we'll consider as the reference point for "forward."
To his left is Ann, to his right is Cathy.
If Bob is facing forward, then:
Ann is positioned to Bob's left.
Cathy is positioned to Bob's right.
Ann and Cathy are facing backwards.
Both Ann and Cathy are oriented in the opposite direction to Bob. If Bob is facing forward, then Ann and Cathy are facing backward.
2. Visualizing the PositionsTo better understand the scenario, let's visualize the positions: Copy
Forward Direction: ↑
Bob (facing forward) | | Ann (facing backward) | / | / | / | / | / | / | / |/ |
And then only the character | in a newline forever.
Re: OpenAI O3-Mini
#459Earlier quoted context omitted.
It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.
Yeah. I immediately thought: I wonder if that 56% is in one or two categories and the rest are worse?
Re: OpenAI O3-Mini
#460I’ll take the China Deluxe instead, actually. I’ve been incredibly pleased with DeepSeek this past week. Wonderful product, I love seeing its brain when it’s thinking.
I recently tried Gemini-1.5-Pro for the first time. It was clearly better than DeepSeek or any of the OpenAI models available to Plus subscribers.