Live data from Hacker News

OpenAI O3-Mini

openai.com

251–260 of 944 posts

Re: OpenAI O3-Mini

#251

Earlier quoted context omitted.

They really need someone in marketing. If the model is for technical stuff, then call it the technical model. How is anyone supposed to know what these model names mean? The only page of theirs attempting to explain this is a total disaster. https://platform.openai.com/docs/models

Yes, this $300Bn company generating +$3.4Bn in revenue needs to hire marketing expert. They can begin by sourcing ideas from us here to save their struggling business from total marketing disaster.

Hype based marketing can be effective but it is high risk and unstable.

A marketing team isn’t a generality that makes a company known, it often focuses on communicating what products different types of customers need from your lineup.

If I sell three medications:

Steve

56285

Priximetrin

And only tell you they are all pain killers but for different types and levels of pain I’m going to leave revenue on the floor. That is no matter how valuable my business is or how well it’s known.

Re: OpenAI O3-Mini

#252
post #20

Earlier quoted context omitted.

There's no moat, and they have to work even harder. Competition is good.

I really don't think this is true. OpenAI has no moat because they have nothing unique; they're using mostly other people's (like Transformers) architectures and other companies hardware. Their value-prop (moat) is that they've burnt more money than everybody else. That moat is trivially circumvented by lighting a larger pile of money and less trivially by lighting the pile more efficently. OpenAI isn't the only comp…

Brand is a moat

Re: OpenAI O3-Mini

#253

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

Yes I'd bet most users just 50/50 it, which actually makes it more remarkable that there was a 56% selection rate

Re: OpenAI O3-Mini

#254

> While OpenAI o1 remains our broader general knowledge reasoning model, OpenAI o3-mini provides a specialized alternative for technical domains requiring precision and speed. I feel like this naming scheme is growing a little tired. o1 is for general knowledge reasoning, o3-mini replaces o1-mini but might be more specialized than o1 for certain technical domains...the "o" in "4o" is for "omni" (referring to its mult…

This is definitely intentional. You can like Sama or dislike him, but he knows how to market a product. Maybe this is a bad call on his part, but it is a call.

That's like making a second reading and appealing to authority.

The naming is bad. Other people already said it you can "google" stuff, you can "deepseek" something, but to "chatgpt" sounds weird.

The model naming is even weirder, like, did they really avoid o2 because of oxigen?

Re: OpenAI O3-Mini

#255

I’ll take the China Deluxe instead, actually. I’ve been incredibly pleased with DeepSeek this past week. Wonderful product, I love seeing its brain when it’s thinking.

Agreed. These locked-down, proprietary models do not interest me. And I certainly am not building product with them - being shackled to a specific provider is a needless business risk.

Re: OpenAI O3-Mini

#256

The naming convention is so messed up. o1, o3-mini (no o2, no o3???)

There's an o1-mini, there's an o3 it just hasn't gone live yet: https://openai.com/12-days/#day-12

they can't call it o2 because: https://en.wikipedia.org/wiki/The_O2_Arena

and the venue's sponsor: https://en.wikipedia.org/wiki/O2_(UK)

Re: OpenAI O3-Mini

#257
post #237

Earlier quoted context omitted.

Maybe they found a need to quantize it further for release, or lobotomise it with more "alignment".

> lobotomise Anyone can write very fast software if you don't mind it sometimes crashing or having weird bugs. Why do people try to meme as if AI is different? It has unexpected outputs sometimes, getting it to not do that is 50% "more alignment" and 50% "hallucinate less". Just today I saw someone get the Amazon bot to roleplay furry erotica. Funny, sure, but it's still obviously a bug that a *sales bot* would do th…

They are implying the release was rushed and they had to reduce the functionality of the model in order to make sure it did not teach people how to make dirty bombs

Re: OpenAI O3-Mini

#258
post #50

It looks like a pretty significant increase on SWE-Bench. Although that makes me wonder if there was some formatting or gotcha that was holding the results back before. If this will work for your use case then it could be a huge discount versus o1. Worth trying again if o1-mini couldn't handle the task before. $4/million output tokens versus $60. https://platform.openai.com/docs/pricing I am Tier 5 but I don't believ…

Genuinely curious, What made you choose OpenAI as your preferred api provider? Its always been the least attractive to me.

We extensively used the batch APIs to decrease cost and handle large amount of data. I also need JSON responses for a lot of things and OpenAI seem to have the best json schema output option out there.

Re: OpenAI O3-Mini

#259
post #37
post #12

Earlier quoted context omitted.

I think OpenAI really needs to rethink its product naming, especially now that they have a portfolio where there's no such clear hierarchy, but they have a place along different axis (speed, cost, reasoning, capabilities, etc). Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.

They're strongly tied to Microsoft, so confusing branding is to be expected.

Flashbacks of the .NET zoo. At least they reigned that in.

Re: OpenAI O3-Mini

#260

Earlier quoted context omitted.

I think this is with and without "tools." They explain it in the system card: > We evaluate SWE-bench in two settings: > *• Agentless*, which is used for all models except o3-mini (tools). This setting uses the Agentless 1.0 scaffold, and models are given 5 tries to generate a candidate patch. We compute pass@1 by averaging the per-instance pass rates of all samples that generated a valid (i.e., non-empty) patch. If…

So am I to understand that they used their internal tooling scaffold on the o3(tools) results only? Because if so, I really don't like that. While it's nonetheless impressive that they scored 61% on SWE-bench with o3-mini combined with their tool scaffolding, comparing Agentless performance with other models seems less impressive, 40% vs 35% when compared to o1-mini if you look at the graph on page 28 of their system…

YC usually says “a startup is the point in your life where tricks stop working”.

Sam Altman is somehow finding this out now, the hard way.

Most paying customers will find out within minutes whether the models can serve their use case, a benchmark isn’t going to change that except for media manipulation (and even that doesn’t work all that well, since journalists don’t really know what they are saying and readers can tell).

Post reply on HN