Live data from Hacker News

OpenAI O3-Mini

openai.com

741–750 of 944 posts

Re: OpenAI O3-Mini

#741

Earlier quoted context omitted.

Like?

Even though it was told that it MUST quote users directly, it still outputs: > It’s already a game changer for many people. But to have so many names like o1, o3-mini, GPT-4o, & GPT-4o-mini suggests there may be too much focus on internal tech details rather than clear communication." (paraphrase based on multiple similar sentiments) It also hallucinates quotes. For example: > "I’m pretty sure 'o3-mini' works better…

In addition to that, it has a section dedicated all to Model Naming and Branding Confusion, but then it puts the following comment in the Performancce and Benchmarking section, even though the value of the comment is ostensibly more to do with the naming being a hindrance rather than make a valuable remark on the benchmarking, which is more of a casualty to the naming confusion:

"The model naming all around is so confusing. Very difficult to tell what breakthrough innovations occurred." – patrickhogan1"

Re: OpenAI O3-Mini

#742
After o3 was announced, with the numbers suggesting it was a major breakthrough, I have to say I’m absolutely not impressed with this version.

I think o1 works significantly better, and that makes me think the timing is more than just a coincidence.

Last week Nvidia lost 600 billion because of DeepSeek R1, and now OpenAI comes out with a new release which feels like it has nothing to do with the promises that were being made about o3.

Re: OpenAI O3-Mini

#743
post #12
post #3

So far, it seems like this is the hierarchy o1 > GPT-4o > o3-mini > o1-mini > GPT-4o-mini o3 mini system card: https://cdn.openai.com/o3-mini-system-card.pdf

I think OpenAI really needs to rethink its product naming, especially now that they have a portfolio where there's no such clear hierarchy, but they have a place along different axis (speed, cost, reasoning, capabilities, etc). Your summary attempt e.g. also misses o3-mini vs o3-mini-high. Lots of trade-ofs.

Did they even think about what happens when they get to o4? We’re going to have GPT-4o and o4

Re: OpenAI O3-Mini

#744
post #669
post #306

I used o3-mini to summarize this thread so far. Here's the result: https://gist.github.com/simonw/09e5922be0cbb85894cf05e6d75ae... For 18,936 input, 2,905 output it cost 3.3612 cents. Here's the script I used to do it: https://til.simonwillison.net/llms/claude-hacker-news-themes...

Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.

I agree with the first part, but I disagree with the conjecture. There are also people who enjoy writing comments, and inasmuch they have to read a bit of context (at least the comment they are replying to). Those will always exits.

Re: OpenAI O3-Mini

#745
post #726

Earlier quoted context omitted.

But extremely rare for a popular wrong comment to not have sub comments debunking them, they are much more reliable than articles therefore.

I’ve found the more I know about the topic at hand, the more wildly many of the comments seem off base, even the highly upvoted undebunked ones. Its harder for me to judge topics I don’t know much about, but I have to assume it’s something similar.

That is true for articles as well though, I find comments typically have better info than the articles. It is more likely for some of the comments to have been written by real experts than that the article is.

Re: OpenAI O3-Mini

#746

Earlier quoted context omitted.

On both HN & Reddit, I find the comments more informative and less frustrating than reading the article usually. But I guess YMMV.

10 years ago Reddit used to be a place where you would get informed opinions and less spam. 5 years ago, HN used to be a place where you would get informed opinions and less spam. Neither of them will go back to the same level of quality. Not anymore.

you need to expend resources (e.g. "proof of work") to post, to drive away low effort spam. https://stacker.news/ is an interesting experiment in that regard.

Re: OpenAI O3-Mini

#747

After o3 was announced, with the numbers suggesting it was a major breakthrough, I have to say I’m absolutely not impressed with this version. I think o1 works significantly better, and that makes me think the timing is more than just a coincidence. Last week Nvidia lost 600 billion because of DeepSeek R1, and now OpenAI comes out with a new release which feels like it has nothing to do with the promises that were be…

Having tried using it, it is much worse than r1. Both the standard and high effort version.

Re: OpenAI O3-Mini

#748
post #682

Earlier quoted context omitted.

Try Google’s NotebookLM

I put one of my own blog posts through NotebookLM soon after it became available, it hallucinated content I didn't write and missed out things I had written. Nice TTS, but otherwise I found it unimpressive.

I’m not talking about the TTS and podcast creation. I’m talking about just asking questions where it gives you the answer with citations.

Re: OpenAI O3-Mini

#750

Earlier quoted context omitted.

You cannot compare GPT-4o and o*(-mini) because GPT-4o is not a reasoning model.

Why can't you ask both questions (on a variety of topics etc), and grade the answers vs an ideal answer? Ends before means. If 4o answered better than o3, would you still use 03 for your task just because you were told it can "reason"?

The point is that you cannot make a general statement that “o1 is better than 4o.”
Post reply on HN