Earlier quoted context omitted.
Have you considered the possibility that your feedback is used to choose what type of response to give to you specifically in the future? I would not consider purposely giving inaccurate feedback for this reason alone.
I don't want a model that's customized to my preferences. My preferences and understanding changes all the time. I want a single source model that's grounded in base truth. I'll let the model know how to structure it in my prompt.
OpenAI O3-Mini
431–440 of 944 posts
Re: OpenAI O3-Mini
#432Re: OpenAI O3-Mini
#433Wow - this is seriously fast (o3-mini), and my initial impressions are very favourable. I was asking it to layout quite a complex html form from a schema and it did a very good job. Looking at the comments on here and the benchmark results I was expecting it to be a bit meh, but initial impressions are quite the opposite I was expecting it to perhaps be a marginal improvement for complex things that need a lot of 're…
It’s 2x the price of R1: https://x.com/deedydas/status/1885440582103031940/photo/1 Is it twice as good though?
Re: OpenAI O3-Mini
#434It looks like a pretty significant increase on SWE-Bench. Although that makes me wonder if there was some formatting or gotcha that was holding the results back before. If this will work for your use case then it could be a huge discount versus o1. Worth trying again if o1-mini couldn't handle the task before. $4/million output tokens versus $60. https://platform.openai.com/docs/pricing I am Tier 5 but I don't believ…
Re: OpenAI O3-Mini
#435Earlier quoted context omitted.
Have you considered the possibility that your feedback is used to choose what type of response to give to you specifically in the future? I would not consider purposely giving inaccurate feedback for this reason alone.
I don't want a model that's customized to my preferences. My preferences and understanding changes all the time. I want a single source model that's grounded in base truth. I'll let the model know how to structure it in my prompt.
Re: OpenAI O3-Mini
#436Earlier quoted context omitted.
Have you considered the possibility that your feedback is used to choose what type of response to give to you specifically in the future? I would not consider purposely giving inaccurate feedback for this reason alone.
I don't want a model that's customized to my preferences. My preferences and understanding changes all the time. I want a single source model that's grounded in base truth. I'll let the model know how to structure it in my prompt.
Re: OpenAI O3-Mini
#437Claude, GPT 4 onwards, and DeepSeek all feel the same to me. Okay to a point, then kinda useless. More like a more convenient specialised Google that you need to double check the results of.
Re: OpenAI O3-Mini
#438Earlier quoted context omitted.
I think OpenAI really needs to rethink its product naming, especially now that they have a portfolio where there's no such clear hierarchy, but they have a place along different axis (speed, cost, reasoning, capabilities, etc). Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.
Can't wait for the eventual rename to GPT Core, GPT Plus, GPT Pro, and GPT Pro Max models! I can see it now: > Unlock our industry leading reasoning features by upgrading to the GPT 4 Pro Max plan.
Re: OpenAI O3-Mini
#439Re: OpenAI O3-Mini
#440Earlier quoted context omitted.
People could be flipping a coin and the score would be the same.
A 12% margin is literally the opposite of a coin flip. Unless you have a really bad coin.
1) Coming off as a jerk, and from a new account is a bad look
2) "Literally the opposite of a coin flip" would probably be either 0% or 100%
3) Your reasoning doesn't stand up without further info; it entirely depends on the sample size. I could have 5 coin flips all come up heads, but over thousands or millions it averages to 50%. 56% on a small sample size is absolutely within margin of error/noise. 56% on a MASSIVE sample size is _statistically_ significant, but isn't even still that much to brag about for something that I feel like they probably intended to be a big step forward.