Live data from Hacker News

OpenAI O3-Mini

openai.com

431–440 of 944 posts

Re: OpenAI O3-Mini

#431

Earlier quoted context omitted.

Have you considered the possibility that your feedback is used to choose what type of response to give to you specifically in the future? I would not consider purposely giving inaccurate feedback for this reason alone.

I don't want a model that's customized to my preferences. My preferences and understanding changes all the time. I want a single source model that's grounded in base truth. I'll let the model know how to structure it in my prompt.

You know there's no such as base truth here? You want to write something like this to start your prompts, "Respond in English, using standard capitalization and punctuation, following rules of grammar as written by Strunk & White, where numbers are represented using arabic numerals in base 10 notation...."???

Re: OpenAI O3-Mini

#433
post #239

Wow - this is seriously fast (o3-mini), and my initial impressions are very favourable. I was asking it to layout quite a complex html form from a schema and it did a very good job. Looking at the comments on here and the benchmark results I was expecting it to be a bit meh, but initial impressions are quite the opposite I was expecting it to perhaps be a marginal improvement for complex things that need a lot of 're…

It’s 2x the price of R1: https://x.com/deedydas/status/1885440582103031940/photo/1 Is it twice as good though?

Whilst I had tried R1 before, I hadn't paid attention to how fast it was. I just tried some similar prompts and was pretty impressed with speed and quality. I think o3-mini was still a bit quicker though.

Re: OpenAI O3-Mini

#434
post #50

It looks like a pretty significant increase on SWE-Bench. Although that makes me wonder if there was some formatting or gotcha that was holding the results back before. If this will work for your use case then it could be a huge discount versus o1. Worth trying again if o1-mini couldn't handle the task before. $4/million output tokens versus $60. https://platform.openai.com/docs/pricing I am Tier 5 but I don't believ…

Tier 5 and I got it almost instantly

Re: OpenAI O3-Mini

#435

Earlier quoted context omitted.

Have you considered the possibility that your feedback is used to choose what type of response to give to you specifically in the future? I would not consider purposely giving inaccurate feedback for this reason alone.

I don't want a model that's customized to my preferences. My preferences and understanding changes all the time. I want a single source model that's grounded in base truth. I'll let the model know how to structure it in my prompt.

A lot of preferences have nothing to do with any truth. Do you like code segments or full code? Do you like paragraphs or bullet points? Heck, do you want English or Japanaese?

Re: OpenAI O3-Mini

#436

Earlier quoted context omitted.

Have you considered the possibility that your feedback is used to choose what type of response to give to you specifically in the future? I would not consider purposely giving inaccurate feedback for this reason alone.

I don't want a model that's customized to my preferences. My preferences and understanding changes all the time. I want a single source model that's grounded in base truth. I'll let the model know how to structure it in my prompt.

What is base truth for e.g. creative writing?

Re: OpenAI O3-Mini

#437
Is AI fizzing out or just me? I feel like they're trying to smash out new models as fast as they can but in reality they're barely any different, it's turning into the smartphone market. New iPhone with a slightly better camera and slightly differently bevelled edges, get it NOW! But doesn't actually do anything better than the iPhone 6.

Claude, GPT 4 onwards, and DeepSeek all feel the same to me. Okay to a point, then kinda useless. More like a more convenient specialised Google that you need to double check the results of.

Re: OpenAI O3-Mini

#438
post #12

Earlier quoted context omitted.

I think OpenAI really needs to rethink its product naming, especially now that they have a portfolio where there's no such clear hierarchy, but they have a place along different axis (speed, cost, reasoning, capabilities, etc). Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.

Can't wait for the eventual rename to GPT Core, GPT Plus, GPT Pro, and GPT Pro Max models! I can see it now: > Unlock our industry leading reasoning features by upgrading to the GPT 4 Pro Max plan.

ngl I'd find that easier to follow lol

Re: OpenAI O3-Mini

#440
post #320

Earlier quoted context omitted.

People could be flipping a coin and the score would be the same.

A 12% margin is literally the opposite of a coin flip. Unless you have a really bad coin.

You're being downvoted for 3 reasons:

1) Coming off as a jerk, and from a new account is a bad look

2) "Literally the opposite of a coin flip" would probably be either 0% or 100%

3) Your reasoning doesn't stand up without further info; it entirely depends on the sample size. I could have 5 coin flips all come up heads, but over thousands or millions it averages to 50%. 56% on a small sample size is absolutely within margin of error/noise. 56% on a MASSIVE sample size is _statistically_ significant, but isn't even still that much to brag about for something that I feel like they probably intended to be a big step forward.

Post reply on HN