Live data from Hacker News

OpenAI O3-Mini

openai.com

491–500 of 944 posts

Re: OpenAI O3-Mini

#491
the o3-mini model would be useful to me if coding's the only thing I need to do in a chat log.

When I use ChatGPT these days, it's to help me write coding videos and then the social media posts around those videos. So that's two specialties in one chat log.

Re: OpenAI O3-Mini

#492

Earlier quoted context omitted.

It's definitely making some errors

Like?

Even though it was told that it MUST quote users directly, it still outputs:

> It’s already a game changer for many people. But to have so many names like o1, o3-mini, GPT-4o, & GPT-4o-mini suggests there may be too much focus on internal tech details rather than clear communication." (paraphrase based on multiple similar sentiments)

It also hallucinates quotes.

For example:

> "I’m pretty sure 'o3-mini' works better for that purpose than 'GPT 4.1.3'." – TeMPOraL

But that comment is not in the user TeMPOraL's comment history.

Sentiment analysis is also faulty.

For example:

> "I’d bet most users just 50/50 it, which actually makes it more remarkable that there was a 56% selection rate." – jackbrookes – This quip injects humor into an otherwise technical discussion about evaluation metrics.

It's not a quip though. That comment was meant in earnest

Re: OpenAI O3-Mini

#493
post #307

Earlier quoted context omitted.

Funny - I had ChatGPT document some stuff for me this week and asked which responses I preferred as well. Didn’t bother reading either of them, just selected one and went on with my day. If it were me I would have set up a “hey do you mind if we give you two results and you can pick your favorite?” prompt to weed out people like me.

I'm surprised how many people claim to do this. You can just not select one.

We -- the people who live in front of a computer -- have been training ourselves to avoid noticing annoyances like captchas, advertising, and GDPR notices for quite a long time.

We find what appears to be the easiest combination "Fuck off, go away" buttons and use them without a moment of actual consideration.

(This doesn't mean that it's actually the easiest method.)

Re: OpenAI O3-Mini

#494

Earlier quoted context omitted.

I like deepseek a lot. But they are currently very glitchy. The API service goes up and down a lot. Maybe they'll sort that out soon.

Apparently they're under a very targeted DDoS for almost a month, with technical details shared in Chinese but very little discussion in English. Which is surprising, it's not like major AI products are getting DDoSed out of existence every day.

Where are the details in Chinese?

Re: OpenAI O3-Mini

#495

Earlier quoted context omitted.

I don't want a model that's customized to my preferences. My preferences and understanding changes all the time. I want a single source model that's grounded in base truth. I'll let the model know how to structure it in my prompt.

You know there's no such as base truth here? You want to write something like this to start your prompts, "Respond in English, using standard capitalization and punctuation, following rules of grammar as written by Strunk & White, where numbers are represented using arabic numerals in base 10 notation...."???

actually, I might appreciate that.

I like precision of language, so maybe just have a system prompt that says "use precise language (ex: no symbolism of any kind)"

Re: OpenAI O3-Mini

#497

Earlier quoted context omitted.

Typically in these tests you have three options "A is better", "B is better" or "they're equal/can't decide". So if 56% prefer O3 Mini, it's likely that way less than half prefer O1.also, the way I understand it, they're comparing a mini model with a large one.

If you use ChatGPT, it sometimes gives you two versions of its response, and you have to choose one or the other if you want to continue prompting. Sure, not picking a response might be a third category. But if that's how they were approaching the analysis, they could have put out a more favorable-looking stat.

> If you use ChatGPT, it sometimes gives you two versions

Does no one else hate it when this happens (especially when on a handheld device)?

Re: OpenAI O3-Mini

#498

Can't wait to try this. What's amazing to me is that when this was revealed just one short month ago, the AI landscape looked very different than it does today with more AI companies jumping into the fray with very compelling models. I wonder how the AI shift has affected this release internally, future releases and their mindset moving forward... How does the efficiency change, the scope of their models, etc.

I thought it was o3 that was released one month ago and received high scores on ARC Prize - https://arcprize.org/blog/oai-o3-pub-breakthrough If they were the same, I would have expected explicit references to o3 in the system card and how o3-mini is distilled or built from o3 - https://cdn.openai.com/o3-mini-system-card.pdf - but there are no references. Excited at the pace all the same. Excited to dig in. The model…

Yeah - the naming is confusing. We're seeing o3-mini. o3 yields marginally better performance given exponentially more compute. Unlike OpenAI, customers will not have an option to throw an endless amount of money at specific tasks/prompts.

Re: OpenAI O3-Mini

#499

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.

It’s 3x cheaper and faster

Re: OpenAI O3-Mini

#500
post #41

Earlier quoted context omitted.

How would the DeepSeek fit into this? Or can it not compare? I don't know much about this stuff, but I've heard recently many people talk about DeepSeek and how unexpected it was.

Deepseek V3 is equivalent to 4o. Deepseek R1 is equivalent to o1 (if not better) I think someone should just build an AI model comparing website at this point. Include all benchmarks and pricing

I had resubscribed to use o1 2 weeks ago and haven't even logged in this week because of R1.

One thing I notice that is huge is being able to see the chain of thought lets me see when my prompt was lacking and the model is a bit confused on what I want.

If I was anymore impressed with R1 I would probably start getting accused of being a CCP shill or wumao lol.

With that said, I think it is very hard to compare models for your own use case. I do suspect there is a shiny new toy bias with all this too.

Poor Sonnet 3.5. I have neglected it so much lately I actually don't know if I have a subscription or not right now.

I do expect an Anthropic reasoning model though to blow everything else away.

Post reply on HN