I really don't get the point of those oX-mini models for chat apps. (API is different, we can benchmark multiple models for a given recurring taks and choose the best one taking costs into consideration). As part of my job, I am trying to promote usage of AI in my company (~150 FTE); we have an OpenAI chatGPT plus subscription for all employees. Roughly speaking the message is: "use GPT-4o all the time, use o1 (soon…
OpenAI O3-Mini
171–180 of 944 posts
Re: OpenAI O3-Mini
#172Earlier quoted context omitted.
There is no such thing. "Meaning" isn't a property of a label, it arises from how that label is used with other labels in communication. It's actually the reason LLMs work in the first place.
You're gonna need to ground those labels in something physical at some point. No one's going to let an LLM near anything important until then.
Re: OpenAI O3-Mini
#173> While OpenAI o1 remains our broader general knowledge reasoning model, OpenAI o3-mini provides a specialized alternative for technical domains requiring precision and speed. I feel like this naming scheme is growing a little tired. o1 is for general knowledge reasoning, o3-mini replaces o1-mini but might be more specialized than o1 for certain technical domains...the "o" in "4o" is for "omni" (referring to its mult…
Re: OpenAI O3-Mini
#174Earlier quoted context omitted.
I did update my comment, but said that I am using the distilled version, so yes?
Even the full model scores below Claude on livebench so a distilled version will likely be even worse.
Re: OpenAI O3-Mini
#175Re: OpenAI O3-Mini
#176Can't wait to try this. What's amazing to me is that when this was revealed just one short month ago, the AI landscape looked very different than it does today with more AI companies jumping into the fray with very compelling models. I wonder how the AI shift has affected this release internally, future releases and their mindset moving forward... How does the efficiency change, the scope of their models, etc.
There's no moat, and they have to work even harder. Competition is good.
Their value-prop (moat) is that they've burnt more money than everybody else. That moat is trivially circumvented by lighting a larger pile of money and less trivially by lighting the pile more efficently.
OpenAI isn't the only company. The Tech companies being beaten massively by Microsoft in #of H100s purchases are the ones with a moat. Google / Amazon with their custom AI chips are going to have a better performance per cost than others and that will be a moat. If you want to get the same performance per cost then you need to spend the time making your own chips which is years of effort (=moat).
Re: OpenAI O3-Mini
#177Earlier quoted context omitted.
Being able to see the thinking trace in R1 is so useful, as you can go back and see if it's getting stuck, making a wrong assumption, missing data, etc. To me that makes it materially more useful than the OpenAI reasoning models, which seem impressive, but are much harder to inspect/debug.
I would actually love if it would just ask me simple questions (just yes/no) when its thinking about something i wasnt clear about and i could help it this way, its a bit sad seeing it write out the assumption and then take the wrong conclusion
Re: OpenAI O3-Mini
#178I’ll take the China Deluxe instead, actually. I’ve been incredibly pleased with DeepSeek this past week. Wonderful product, I love seeing its brain when it’s thinking.
I am running the 7B distilled version locally. I asked it to create a skeleton MEAN project. Everything was great but then it started to generate the front-end and I noticed the file extension (.tsx) and then saw react getting imported. I gave the same prompt to sonnet 3.5 and not a single hiccup. Maybe not an indication that Deepseek is worse/bad (I am using a distilled version), but moreso speaks to much react/next…
Re: OpenAI O3-Mini
#179Earlier quoted context omitted.
It's like AWS SKU naming (`c5d.metal`, `p5.48xlarge`, etc.), except non-technical consumers are expected to understand it.
Have you seen Azure VM SKU naming? It's.. impressive.
Re: OpenAI O3-Mini
#180I understand that keeping the same data and curating it might be beneficial. But it sounds odd to roll back in time with the knowledge cutoff. AFAIK, the only event that happened around that time was the start of the Gaza conflict.