Live data from Hacker News

OpenAI O3-Mini

openai.com

291–300 of 944 posts

Re: OpenAI O3-Mini

#291

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

Funny - I had ChatGPT document some stuff for me this week and asked which responses I preferred as well.

Didn’t bother reading either of them, just selected one and went on with my day.

If it were me I would have set up a “hey do you mind if we give you two results and you can pick your favorite?” prompt to weed out people like me.

Re: OpenAI O3-Mini

#293

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

I almost always pick the second one, because it's closer to the submit button and the one I read first.

Re: OpenAI O3-Mini

#294

This took 1:53 in o3-mini https://chatgpt.com/share/679d310d-6064-8010-ba78-6bd5ed3360... The 4o model without using the Python tool https://chatgpt.com/share/679d32bd-9ba8-8010-8f75-2f26a792e0... Trying to get accurate results with the paid version of 4o with the Python interpreter. https://chatgpt.com/share/679d31f3-21d4-8010-9932-7ecadd0b87... The share link doesn’t show the output for some reason. But it did work…

I would not expect any LLM to get this right. I think people have too high expectations for it.

Now if you asked it to write a Python program to list them in order, and have it enter all the names, birthdays, and year elected in a list to get the program to run - that's more reasonable.

Re: OpenAI O3-Mini

#295

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

Those prompts are so irritating and so frequent that I’ve taken to just quickly picking whichever one looks worse at a cursory glance. I’m paying them, they shouldn’t expect high quality work from me.

That's such a counter-productive and frankly dumb thing to do. Just don't vote on them.

Re: OpenAI O3-Mini

#296
post #267

Earlier quoted context omitted.

I really don't think this is true. OpenAI has no moat because they have nothing unique; they're using mostly other people's (like Transformers) architectures and other companies hardware. Their value-prop (moat) is that they've burnt more money than everybody else. That moat is trivially circumvented by lighting a larger pile of money and less trivially by lighting the pile more efficently. OpenAI isn't the only comp…

> That moat is trivially circumvented by lighting a larger pile of money and less trivially by lighting the pile more efficently. Google with all its money and smart engineers was not able to build a simple chat application.

But with their internal progression structure they can build and cancel eight mediocre chat apps.

Re: OpenAI O3-Mini

#298

Earlier quoted context omitted.

Those prompts are so irritating and so frequent that I’ve taken to just quickly picking whichever one looks worse at a cursory glance. I’m paying them, they shouldn’t expect high quality work from me.

Have you considered the possibility that your feedback is used to choose what type of response to give to you specifically in the future? I would not consider purposely giving inaccurate feedback for this reason alone.

Alternatively, I'll use the tool that is most user friendly and provides the most value for my money.

Wasting time on an anti pattern is not value nor is it trying to outguess the way that selection mechanism is used.

Re: OpenAI O3-Mini

#299
post #273

Earlier quoted context omitted.

That's like making a second reading and appealing to authority. The naming is bad. Other people already said it you can "google" stuff, you can "deepseek" something, but to "chatgpt" sounds weird. The model naming is even weirder, like, did they really avoid o2 because of oxigen?

> but to "chatgpt" sounds weird. People just say it differently, they say "ask chatgpt"

I normally use Claude, but "Ask Claude", but unless it's someone who knows me well, I say "Ask ChatGPT", or it's just not as claer; and I don't think it's primarily due to popularity.

Re: OpenAI O3-Mini

#300
post #208

Earlier quoted context omitted.

Genuinely curious, What made you choose OpenAI as your preferred api provider? Its always been the least attractive to me.

Who else might be a good choice? Deepseek is down. Who has the cheapest gpt3.5 level or above api

Run it locally, the distilled smaller ones aren't bad at all.
Post reply on HN