I ran some quick programming tasks I have used O1 previously: 1. 1/4th time for reasoning for most tasks. 2. Far better results.
Compared to o1 or o1-pro?
OpenAI O3-Mini
311–320 of 944 posts
Re: OpenAI O3-Mini
#312Earlier quoted context omitted.
Those prompts are so irritating and so frequent that I’ve taken to just quickly picking whichever one looks worse at a cursory glance. I’m paying them, they shouldn’t expect high quality work from me.
Have you considered the possibility that your feedback is used to choose what type of response to give to you specifically in the future? I would not consider purposely giving inaccurate feedback for this reason alone.
I want a single source model that's grounded in base truth. I'll let the model know how to structure it in my prompt.
Re: OpenAI O3-Mini
#313Earlier quoted context omitted.
Deepseek is the state of the art right now in terms of performance and output. It's really fast. The way it "explains" how it's thinking is remarkable.
DeepSeek is great because: 1) you can run the model locally, 2) the research was openly shared, and 3) the reasoning tokens are open. It is not, in my experience, state of the art. In all of my side by side comparisons thus far in real world applications between DeepSeek V3 and R1 vs 4o and o1, the latter has always performed better. OpenAI's models are also more consistent, glitching out maybe one in 10,000, whereas…
Are you running the smaller models locally? Doesn't seems unfair to compare it against 4o and o1 behind OpenAI APIs.
Re: OpenAI O3-Mini
#314This took 1:53 in o3-mini https://chatgpt.com/share/679d310d-6064-8010-ba78-6bd5ed3360... The 4o model without using the Python tool https://chatgpt.com/share/679d32bd-9ba8-8010-8f75-2f26a792e0... Trying to get accurate results with the paid version of 4o with the Python interpreter. https://chatgpt.com/share/679d31f3-21d4-8010-9932-7ecadd0b87... The share link doesn’t show the output for some reason. But it did work…
I would not expect any LLM to get this right. I think people have too high expectations for it. Now if you asked it to write a Python program to list them in order, and have it enter all the names, birthdays, and year elected in a list to get the program to run - that's more reasonable.
DeepSeek also gets the order right.
It doesn’t show on the share link. But it actually outputs the list correctly from the built in Python interpreter.
For some things, ChatGPT 4o will automatically use its Python runtime
Re: OpenAI O3-Mini
#315Earlier quoted context omitted.
I really don't think this is true. OpenAI has no moat because they have nothing unique; they're using mostly other people's (like Transformers) architectures and other companies hardware. Their value-prop (moat) is that they've burnt more money than everybody else. That moat is trivially circumvented by lighting a larger pile of money and less trivially by lighting the pile more efficently. OpenAI isn't the only comp…
> That moat is trivially circumvented by lighting a larger pile of money and less trivially by lighting the pile more efficently. Google with all its money and smart engineers was not able to build a simple chat application.
Re: OpenAI O3-Mini
#316Earlier quoted context omitted.
Those prompts are so irritating and so frequent that I’ve taken to just quickly picking whichever one looks worse at a cursory glance. I’m paying them, they shouldn’t expect high quality work from me.
That's such a counter-productive and frankly dumb thing to do. Just don't vote on them.
Re: OpenAI O3-Mini
#317Earlier quoted context omitted.
They mention in the model card, it's so that they can have a separate "system" role that the user can't change, and they trained the model to prioritise it over the "developer" role, to combat "jailbreaks". Thank God for DeepSeek.
They should have just created something above system and left as it was.
Re: OpenAI O3-Mini
#318Earlier quoted context omitted.
I think OpenAI really needs to rethink its product naming, especially now that they have a portfolio where there's no such clear hierarchy, but they have a place along different axis (speed, cost, reasoning, capabilities, etc). Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.
Can't wait for the eventual rename to GPT Core, GPT Plus, GPT Pro, and GPT Pro Max models! I can see it now: > Unlock our industry leading reasoning features by upgrading to the GPT 4 Pro Max plan.
Re: OpenAI O3-Mini
#319Earlier quoted context omitted.
Those prompts are so irritating and so frequent that I’ve taken to just quickly picking whichever one looks worse at a cursory glance. I’m paying them, they shouldn’t expect high quality work from me.
Have you considered the possibility that your feedback is used to choose what type of response to give to you specifically in the future? I would not consider purposely giving inaccurate feedback for this reason alone.
Re: OpenAI O3-Mini
#320> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…