Earlier quoted context omitted.
So am I to understand that they used their internal tooling scaffold on the o3(tools) results only? Because if so, I really don't like that. While it's nonetheless impressive that they scored 61% on SWE-bench with o3-mini combined with their tool scaffolding, comparing Agentless performance with other models seems less impressive, 40% vs 35% when compared to o1-mini if you look at the graph on page 28 of their system…
YC usually says “a startup is the point in your life where tricks stop working”. Sam Altman is somehow finding this out now, the hard way. Most paying customers will find out within minutes whether the models can serve their use case, a benchmark isn’t going to change that except for media manipulation (and even that doesn’t work all that well, since journalists don’t really know what they are saying and readers can…
OpenAI O3-Mini
671–680 of 944 posts
Re: OpenAI O3-Mini
#672I used o3-mini to summarize this thread so far. Here's the result: https://gist.github.com/simonw/09e5922be0cbb85894cf05e6d75ae... For 18,936 input, 2,905 output it cost 3.3612 cents. Here's the script I used to do it: https://til.simonwillison.net/llms/claude-hacker-news-themes...
Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.
Re: OpenAI O3-Mini
#673Sure as a clock, tick follows tock. Can't imagine trying to build out cost structures, business plans, product launches etc on such rapidly shifting sands. Good that you get more for your money, I suppose. But I get the feeling no model or provider is worth committing to in any serious way.
Re: OpenAI O3-Mini
#674So far, it seems like this is the hierarchy o1 > GPT-4o > o3-mini > o1-mini > GPT-4o-mini o3 mini system card: https://cdn.openai.com/o3-mini-system-card.pdf
You cannot compare GPT-4o and o*(-mini) because GPT-4o is not a reasoning model.
Ends before means.
If 4o answered better than o3, would you still use 03 for your task just because you were told it can "reason"?
Re: OpenAI O3-Mini
#675Earlier quoted context omitted.
Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.
That would be an actual improvement. Reading the comments section usually just leads to personal energy waste.
Re: OpenAI O3-Mini
#676Earlier quoted context omitted.
Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.
That would be an actual improvement. Reading the comments section usually just leads to personal energy waste.
Re: OpenAI O3-Mini
#677Earlier quoted context omitted.
That would be an actual improvement. Reading the comments section usually just leads to personal energy waste.
On both HN & Reddit, I find the comments more informative and less frustrating than reading the article usually. But I guess YMMV.
Re: OpenAI O3-Mini
#678Re: OpenAI O3-Mini
#679Earlier quoted context omitted.
That would be an actual improvement. Reading the comments section usually just leads to personal energy waste.
On both HN & Reddit, I find the comments more informative and less frustrating than reading the article usually. But I guess YMMV.
Especially if somebody is being wrong.
Re: OpenAI O3-Mini
#680I used o3-mini to summarize this thread so far. Here's the result: https://gist.github.com/simonw/09e5922be0cbb85894cf05e6d75ae... For 18,936 input, 2,905 output it cost 3.3612 cents. Here's the script I used to do it: https://til.simonwillison.net/llms/claude-hacker-news-themes...
I don't have much experience with reasoning models yet. That's why.