Live data from Hacker News

OpenAI O3-Mini

openai.com

441–450 of 944 posts

Re: OpenAI O3-Mini

#441
Just tested two complicated coding tasks, and surprisingly o3-mini-high nailed it while Sonnet 3.5 failed it. Will do more tests tomorrow.

Re: OpenAI O3-Mini

#442

Earlier quoted context omitted.

I really don't think this is true. OpenAI has no moat because they have nothing unique; they're using mostly other people's (like Transformers) architectures and other companies hardware. Their value-prop (moat) is that they've burnt more money than everybody else. That moat is trivially circumvented by lighting a larger pile of money and less trivially by lighting the pile more efficently. OpenAI isn't the only comp…

Brand is a moat

Their brand is as tainted as Meta's, which was bad enough to merit a rebranding from Facebook.

Re: OpenAI O3-Mini

#443

Earlier quoted context omitted.

"ChatGPT" (chatgpt-4o) is now its own model, distinct from gpt-4o. As for self-limiting usage by non-power users, they're already doing that: ChatGPT app automatically picks a model depending on what capabilities you invoke. While they provide a limited ability to see and switch the model in use, they're clearly expecting regular users not to care, and design their app around that.

None of that matters to normal users, and you could satisfy power users with serial numbers or even unique ideograms. Naming isn't that hard, and their models are surprisingly adept at it. A consistent naming scheme improves customer experience by preventing confusion - when a new model comes out, I field questions for days from friends and family - "what does this mean? which model should i use? Aww, I have to downl…

I agree with your observations, and that they both could and should do better. However, they have the privilege of being the AI company, the most hyped-up brand in the most hyped-up segment of economy - at this point, the impact of their naming strategy is approximately nil. Sure, they're confusing their users a bit, but their users are very highly motivated.

It's like with videogames - most of them commit all kinds of UI/UX sins, and I often wish they didn't, but excepting extreme cases, the players are too motivated to care or notice.

Re: OpenAI O3-Mini

#444
post #437

Is AI fizzing out or just me? I feel like they're trying to smash out new models as fast as they can but in reality they're barely any different, it's turning into the smartphone market. New iPhone with a slightly better camera and slightly differently bevelled edges, get it NOW! But doesn't actually do anything better than the iPhone 6. Claude, GPT 4 onwards, and DeepSeek all feel the same to me. Okay to a point, th…

on the contrary, it's accelerating since they unlocked a new paradigm of scaling

Re: OpenAI O3-Mini

#445

Earlier quoted context omitted.

Running it locally lets you INTERJECT IN IT'S THINKING IN REALTIME and I cannot stress enough how useful that is.

How are you running it locally??

I am running a 4bit imatrix quant of the 70b distill with quantized context. It fits in the 43gb of vram I have.

Re: OpenAI O3-Mini

#446

Earlier quoted context omitted.

Fundamentally the UI is up to you, I have a "typing-pauses-inference-and-starts-gaslighting" feature in my homebrew frontend, but in OpenWebUI/Sillytavern you can just pause it and edit the chain of thought and then have it continue from the edit.

That's a great idea. In your frontend, do you write in the same text entry field as the bot? I use oobabooga/text-generation-webui and I findit's a little awkward to edit the bot responses.

No, but the chat divs are all contenteditable.

Re: OpenAI O3-Mini

#447
post #437

Is AI fizzing out or just me? I feel like they're trying to smash out new models as fast as they can but in reality they're barely any different, it's turning into the smartphone market. New iPhone with a slightly better camera and slightly differently bevelled edges, get it NOW! But doesn't actually do anything better than the iPhone 6. Claude, GPT 4 onwards, and DeepSeek all feel the same to me. Okay to a point, th…

Boiling frog. The advances are happening so rapidly, but incrementally, that it's not being registered. It just seems like the normal state.

Compare LLMs from a year or two ago with the ones out today on practically any task. It's night and day difference.

This is specially so when you start taking into account these "reasoning" models. It's mind blowing how much better they are than "non-reasoning" models for tasks like planning and coding.

https://aider.chat/docs/leaderboards/#aider-polyglot-benchma...

Re: OpenAI O3-Mini

#448

I hope chatgpt reconsiders the naming of their models some time. I have troubles deciding which model is the one I should use.

They release models too often for a new one to be better at everything, so you have to pick the right one for your task.

Re: OpenAI O3-Mini

#449
post #12
post #3

So far, it seems like this is the hierarchy o1 > GPT-4o > o3-mini > o1-mini > GPT-4o-mini o3 mini system card: https://cdn.openai.com/o3-mini-system-card.pdf

I think OpenAI really needs to rethink its product naming, especially now that they have a portfolio where there's no such clear hierarchy, but they have a place along different axis (speed, cost, reasoning, capabilities, etc). Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.

Careful what you wish for. Next thing you know they're going to have names like Betsy and be full of unique quirky behavior to help remind us that they're different people.
Post reply on HN