Live data from Hacker News

OpenAI O3-Mini

openai.com

301–310 of 944 posts

Re: OpenAI O3-Mini

#301

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.

exactly I was surprised as well

Re: OpenAI O3-Mini

#302
post #3

So far, it seems like this is the hierarchy o1 > GPT-4o > o3-mini > o1-mini > GPT-4o-mini o3 mini system card: https://cdn.openai.com/o3-mini-system-card.pdf

I actually switched back from o1-preview to GPT-4o due to tooling integration and web search. I find that more often than not, the ability of GPT-4o to use these tools outweighs o1's improved accuracy.

Re: OpenAI O3-Mini

#303
post #3

So far, it seems like this is the hierarchy o1 > GPT-4o > o3-mini > o1-mini > GPT-4o-mini o3 mini system card: https://cdn.openai.com/o3-mini-system-card.pdf

You cannot compare GPT-4o and o*(-mini) because GPT-4o is not a reasoning model.

Sure you can. "Reasoning" is ultimately an implementation detail, and the only thing that matters for capabilities is results, not process.

Re: OpenAI O3-Mini

#304

Earlier quoted context omitted.

By default, we do not train on any inputs or outputs from our products for business users, including ChatGPT Team, ChatGPT Enterprise, and the API. We offer API customers a way to opt-in to share data with us, such as by providing feedback in the Playground, which we then use to improve our models. Unless they explicitly opt-in, organizations are opted out of data-sharing by default. The business bit is confusing, I…

So for posterity, in this subthread we found that OpenAI indeed trains on user data and it isn't something that only DeepSeek does.

So for posterity, in this subthread we found that I can use OpenAI without them training on my data, whereas I cannot with DeepSeek.

Re: OpenAI O3-Mini

#305

Earlier quoted context omitted.

So for posterity, in this subthread we found that OpenAI indeed trains on user data and it isn't something that only DeepSeek does.

So for posterity, in this subthread we found that I can use OpenAI without them training on my data, whereas I cannot with DeepSeek.

What do you mean? They both say the same thing for usage through API. You can also use DeepSeek on your own compute.

Re: OpenAI O3-Mini

#307

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

Funny - I had ChatGPT document some stuff for me this week and asked which responses I preferred as well. Didn’t bother reading either of them, just selected one and went on with my day. If it were me I would have set up a “hey do you mind if we give you two results and you can pick your favorite?” prompt to weed out people like me.

I'm surprised how many people claim to do this. You can just not select one.

Re: OpenAI O3-Mini

#308
A random idea - train one of those models on you, keep it aside, let it somehow work out your intricacies, moods, details, childhood memories, personality, flaws, strengths. Methods can be various - initial dump of social networks, personal photos and videos, maybe some intense conversation to grok rough you, then polish over time.

A first step to digital immortality, could be a nice startup of some personalized product for rich, and then even regular folks. Immortality not in ourselves as meat bags of course, we die regardless, but digital copy and memento that our children can use if feeling lonely and can carry with themselves anywhere, or later descendants out of curiosity to hold massive events like weddings. One could 'invite' long lost ancestors. Maybe your grand-grand father would be a cool guy you could easily click with these days via verbal input. Heck even 3D detailed model.

An additional service, 'perpetually' paid - keeping your data model safe, taking care of it, backups, heck even maybe give it a bit of computing power to to receive current news in some light fashion and evolve, could be extras. Different tiers for different level of services and care.

Or am I decade or two ahead? I can see this as universally interesting across many if not all cultures.

Re: OpenAI O3-Mini

#309

This took 1:53 in o3-mini https://chatgpt.com/share/679d310d-6064-8010-ba78-6bd5ed3360... The 4o model without using the Python tool https://chatgpt.com/share/679d32bd-9ba8-8010-8f75-2f26a792e0... Trying to get accurate results with the paid version of 4o with the Python interpreter. https://chatgpt.com/share/679d31f3-21d4-8010-9932-7ecadd0b87... The share link doesn’t show the output for some reason. But it did work…

The 4o model's output is blatantly wrong. I'm not going to look up if it's the order or the ages that are incorrect, but: 36. Abraham Lincoln – 52 years, 20 days (1861) 37. James Garfield – 49 years, 105 days (1881) 38. Lyndon B. Johnson – 55 years, 87 days (1963) Basically everything after #15 in the list is scrambled.

That was the point. The 4o model without using Python was wrong. The o3 model worked correctly without needing an external tool

Re: OpenAI O3-Mini

#310

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.

That would be 12%, why would you assume that is eaten by statistical noise?
Post reply on HN