Live data from Hacker News

OpenAI O3-Mini

openai.com

201–210 of 944 posts

Re: OpenAI O3-Mini

#201
post #196

I think that OpenAI should reduce the prices even further to be competitive with Qwen or Deepseek. There are a lot of vendors offering Deepseek R1 for $2-2.5 per 1 million tokens output.

Would you have specific recommendations of such vendors?

Re: OpenAI O3-Mini

#203
post #67
post #35

Earlier quoted context omitted.

I don't think OpenAI is training on your data. At least they say they don't, and I believe that. I wouldn't be surprised if the NSA or something has access to data if they request it or something though. But DeepSeek clearly states in their terms of service that they can train on your API data or use it for other purposes. Which one might assume their government can access as well. We need direct eval comparisons bet…

OpenAI clearly states that they train on your data https://help.openai.com/en/articles/5722486-how-your-data-is...

> Services for businesses, such as ChatGPT Team, ChatGPT Enterprise, and our API Platform > By default, we do not train on any inputs or outputs from our products for business users, including ChatGPT Team, ChatGPT Enterprise, and the API.

So on API they don't train by default, for other paid subscription they mention you can opt-out

Re: OpenAI O3-Mini

#204
post #90

200k context window $1.1/m for input $4.4/m for output I assume thinking medium and hard would consume more tokens. I feel the timing is bad for this release especially when deepseek R1 is still peaking. People will compare and might get disappointed with this model.

The model looks quite a bit better in the benchmarks so unless they overfit the model on them it would probably perform better than deepseek.

Re: OpenAI O3-Mini

#205

Did anyone else notice that o3-mini's SWE bench dropped from 61% in the leaked System Card earlier today to 49.3% in this blog post, which puts o3-mini back in line with Claude on real-world coding tasks? Am I missing something?

I think this is with and without "tools." They explain it in the system card: > We evaluate SWE-bench in two settings: > *• Agentless*, which is used for all models except o3-mini (tools). This setting uses the Agentless 1.0 scaffold, and models are given 5 tries to generate a candidate patch. We compute pass@1 by averaging the per-instance pass rates of all samples that generated a valid (i.e., non-empty) patch. If…

So am I to understand that they used their internal tooling scaffold on the o3(tools) results only? Because if so, I really don't like that.

While it's nonetheless impressive that they scored 61% on SWE-bench with o3-mini combined with their tool scaffolding, comparing Agentless performance with other models seems less impressive, 40% vs 35% when compared to o1-mini if you look at the graph on page 28 of their system card pdf (https://cdn.openai.com/o3-mini-system-card.pdf).

It just feels like data manipulation to suggest that o3-mini is much more performant than past models. A fairer picture would still paint a performance improvement, but it look less exciting and more incremental.

Of course the real improvement is cost, but still, it kind of rubs me the wrong way.

Re: OpenAI O3-Mini

#206
post #20

Can't wait to try this. What's amazing to me is that when this was revealed just one short month ago, the AI landscape looked very different than it does today with more AI companies jumping into the fray with very compelling models. I wonder how the AI shift has affected this release internally, future releases and their mindset moving forward... How does the efficiency change, the scope of their models, etc.

There's no moat, and they have to work even harder. Competition is good.

Capex was the theoretical moat, same as TSMC and similar businesses. DeepSeek poked a hole in this theory. OpenAI will need to deliver massive improvements to justify a 1 billion dollar training cost relative to 5 million dollars.

Re: OpenAI O3-Mini

#207
post #128

Earlier quoted context omitted.

Not really. They’re successful because they created one of the most interesting products in human history, not because they have any idea how to brand it.

If that were the case, they’d be neck and neck with Anthropic and Claude. But ChatGPT has far more market share and name recognition, especially among normies. Branding clearly plays a huge role.

I prefer Anthropic's models but ChatGPT (the web interface) is far superior to Claude IMHO. Web search, long-term memory, and chat history sharing are hard to give up.

Re: OpenAI O3-Mini

#208
post #50

It looks like a pretty significant increase on SWE-Bench. Although that makes me wonder if there was some formatting or gotcha that was holding the results back before. If this will work for your use case then it could be a huge discount versus o1. Worth trying again if o1-mini couldn't handle the task before. $4/million output tokens versus $60. https://platform.openai.com/docs/pricing I am Tier 5 but I don't believ…

Genuinely curious, What made you choose OpenAI as your preferred api provider? Its always been the least attractive to me.

Who else might be a good choice? Deepseek is down. Who has the cheapest gpt3.5 level or above api

Re: OpenAI O3-Mini

#209
They made a discount; it's very impressive; they probably found a very efficient way, so it's discounted. I guess there's no need to build a very large nuclear power plant or a $9 trillion chip factory to run a single large language model. Efficiency has skyrocketed, or thanks to competition, OpenAI's all problems were solved.

Re: OpenAI O3-Mini

#210

Earlier quoted context omitted.

> I wonder if we'll end up with both a 4o and o4... The perplexing thing is that someone has to have said that, right? It has to have been brought up in some meeting when they were brainstorming names that if you have 4o and o1 with the intention of incrementing o1 you'll eventually end up with an o4. Where they really went off the rails was not just bailing when they realized they couldn't use o2. In that moment the…

The obvious solution could be to just keep skipping the even numbers and go to o5.

Or further the hype and name it o9.
Post reply on HN