Live data from Hacker News

OpenAI O3-Mini

openai.com

671–680 of 944 posts

Re: OpenAI O3-Mini

#671

Earlier quoted context omitted.

So am I to understand that they used their internal tooling scaffold on the o3(tools) results only? Because if so, I really don't like that. While it's nonetheless impressive that they scored 61% on SWE-bench with o3-mini combined with their tool scaffolding, comparing Agentless performance with other models seems less impressive, 40% vs 35% when compared to o1-mini if you look at the graph on page 28 of their system…

YC usually says “a startup is the point in your life where tricks stop working”. Sam Altman is somehow finding this out now, the hard way. Most paying customers will find out within minutes whether the models can serve their use case, a benchmark isn’t going to change that except for media manipulation (and even that doesn’t work all that well, since journalists don’t really know what they are saying and readers can…

My guess is this cheap mini-model comes out now after DeepSeek very recently shook the stock-market greatly with its cheap price and relatively good performance. .

Re: OpenAI O3-Mini

#672
post #669
post #306

I used o3-mini to summarize this thread so far. Here's the result: https://gist.github.com/simonw/09e5922be0cbb85894cf05e6d75ae... For 18,936 input, 2,905 output it cost 3.3612 cents. Here's the script I used to do it: https://til.simonwillison.net/llms/claude-hacker-news-themes...

Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.

That would be an actual improvement. Reading the comments section usually just leads to personal energy waste.

Re: OpenAI O3-Mini

#673

Sure as a clock, tick follows tock. Can't imagine trying to build out cost structures, business plans, product launches etc on such rapidly shifting sands. Good that you get more for your money, I suppose. But I get the feeling no model or provider is worth committing to in any serious way.

Terrible time to open a shovel store, amazing time to pick up a shovel.

Re: OpenAI O3-Mini

#674
post #3

So far, it seems like this is the hierarchy o1 > GPT-4o > o3-mini > o1-mini > GPT-4o-mini o3 mini system card: https://cdn.openai.com/o3-mini-system-card.pdf

You cannot compare GPT-4o and o*(-mini) because GPT-4o is not a reasoning model.

Why can't you ask both questions (on a variety of topics etc), and grade the answers vs an ideal answer?

Ends before means.

If 4o answered better than o3, would you still use 03 for your task just because you were told it can "reason"?

Re: OpenAI O3-Mini

#675
post #669

Earlier quoted context omitted.

Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.

That would be an actual improvement. Reading the comments section usually just leads to personal energy waste.

Yes, but who will comment then and on what?

Re: OpenAI O3-Mini

#676
post #669

Earlier quoted context omitted.

Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.

That would be an actual improvement. Reading the comments section usually just leads to personal energy waste.

On both HN & Reddit, I find the comments more informative and less frustrating than reading the article usually. But I guess YMMV.

Re: OpenAI O3-Mini

#677

Earlier quoted context omitted.

That would be an actual improvement. Reading the comments section usually just leads to personal energy waste.

On both HN & Reddit, I find the comments more informative and less frustrating than reading the article usually. But I guess YMMV.

You need to read the article first to know that. But most people won't.

Re: OpenAI O3-Mini

#678
post #677

Earlier quoted context omitted.

On both HN & Reddit, I find the comments more informative and less frustrating than reading the article usually. But I guess YMMV.

You need to read the article first to know that. But most people won't.

You can read the article after the comments.

Re: OpenAI O3-Mini

#679

Earlier quoted context omitted.

That would be an actual improvement. Reading the comments section usually just leads to personal energy waste.

On both HN & Reddit, I find the comments more informative and less frustrating than reading the article usually. But I guess YMMV.

I agree, they are! But reading through them, or even worse, engaging with them, is a serious energy drain.

Especially if somebody is being wrong.

Re: OpenAI O3-Mini

#680
post #306

I used o3-mini to summarize this thread so far. Here's the result: https://gist.github.com/simonw/09e5922be0cbb85894cf05e6d75ae... For 18,936 input, 2,905 output it cost 3.3612 cents. Here's the script I used to do it: https://til.simonwillison.net/llms/claude-hacker-news-themes...

Why use a reasoning model for a summarisation task? Serious question, would it benefit?

I don't have much experience with reasoning models yet. That's why.

Post reply on HN