Live data from Hacker News

OpenAI O3-Mini

openai.com

191–200 of 944 posts

Re: OpenAI O3-Mini

#191

Wake me up when the full o3 is out.

My guess is it will happen right after Sam Altman's next public freakout about how dangerous this new model they have in store is and how it tried to escape from its confinement and kidnap the alignment operator.

Re: OpenAI O3-Mini

#192
post #50

It looks like a pretty significant increase on SWE-Bench. Although that makes me wonder if there was some formatting or gotcha that was holding the results back before. If this will work for your use case then it could be a huge discount versus o1. Worth trying again if o1-mini couldn't handle the task before. $4/million output tokens versus $60. https://platform.openai.com/docs/pricing I am Tier 5 but I don't believ…

Genuinely curious, What made you choose OpenAI as your preferred api provider? Its always been the least attractive to me.

I have mainly been using Claude 3.5/3.6 Sonnet via API in the last several months (or since 3.5 Sonnet came out). However, I was using o1 for a challenging task at one point, but last I tested it had issues with some extra backslashes for that application.

I also have tested with DeepSeek R1 and will test some more with that although in a way Claude 3.6 with CoT is pretty good. Last time I tried to test R1 their API was out.

Re: OpenAI O3-Mini

#193
post #23

why should anyone use this when deepseek is free/cheaper? openai is no longer relevant.

> openai is no longer relevant.

I think you've spent a little too long hitting on the Deepseek pipe. Enterprise customers with familiarity with China will avoid the hosted model for data security and IP protection reasons, among others.

Those working in any area considered economically competitive with China will also be hesitant to use the vanilla model in self-hosted form as there perpetually remains the standing question on what all they've tuned inside the model to benefit the CCP. Perhaps even in subtle ways reminiscent of the Trisolaran sophons from the Three Body Problem.

For instance, you can imagine that if Germany had released an OS model in 1943, that the Americans wouldn't have trusted it to help them develop better military systems even if initial testing passed muster.

Unfortunately, state control of private enterprise in the Chinese economy makes it unproductive to separate the two from one another. Particularly in Deepseek's case as a wide array of Chinese state-linked social media accounts were promoting V3/R1 on the day of its public release.

https://www.reuters.com/technology/artificial-intelligence/c...

Re: OpenAI O3-Mini

#194

Does anyone know why GPT4 has knowledge cutoff December 2023 and all the other models (newer ones like 4o, O1, O3) seem to have knowledge cutoff October 2023? https://platform.openai.com/docs/models#o3-mini I understand that keeping the same data and curating it might be beneficial. But it sounds odd to roll back in time with the knowledge cutoff. AFAIK, the only event that happened around that time was the start of…

I think trained knowledge is less and less important - as these multi-modal models have the ability to search the web and have much larger context windows.

Re: OpenAI O3-Mini

#195
post #3

So far, it seems like this is the hierarchy o1 > GPT-4o > o3-mini > o1-mini > GPT-4o-mini o3 mini system card: https://cdn.openai.com/o3-mini-system-card.pdf

at least if i ran the company you'd know that

ChatGPTasdhjf-final-final-use_this_one.pt > ChatGPTasdhjf-final.pt > ChatGPTasdhjf.pt > ChatGPTasd.pt> ChatGPT.pt

Re: OpenAI O3-Mini

#196
I think that OpenAI should reduce the prices even further to be competitive with Qwen or Deepseek. There are a lot of vendors offering Deepseek R1 for $2-2.5 per 1 million tokens output.

Re: OpenAI O3-Mini

#197
post #35

Earlier quoted context omitted.

I don't think OpenAI is training on your data. At least they say they don't, and I believe that. I wouldn't be surprised if the NSA or something has access to data if they request it or something though. But DeepSeek clearly states in their terms of service that they can train on your API data or use it for other purposes. Which one might assume their government can access as well. We need direct eval comparisons bet…

Yes but DeepSeek models can be accessed through the APIs of Cloudflare or GitHub, in which case no training on your data takes place.

True.

Re: OpenAI O3-Mini

#198
post #171
post #96

I really don't get the point of those oX-mini models for chat apps. (API is different, we can benchmark multiple models for a given recurring taks and choose the best one taking costs into consideration). As part of my job, I am trying to promote usage of AI in my company (~150 FTE); we have an OpenAI chatGPT plus subscription for all employees. Roughly speaking the message is: "use GPT-4o all the time, use o1 (soon…

Why not promote o1? 4o is rather sloppy in comparison

99% of what people use ChatGPT is for very mundane stuff. Think “translate this email to English”, “fix spelling mistakes”, “write this better for me”. Data extraction (list of emails) is big as well. You don’t need o1 for that; and people make lot of those requests per day.

Additionally, o1 does not have access to search and multimodality and taking a screenshot of something and asking questions about it is also a big use case.

It’s easy to overlook how widely ChatGPT is used for very small stuff. But compounded it’s still a game changer for many people.

Re: OpenAI O3-Mini

#200

Does anyone know why GPT4 has knowledge cutoff December 2023 and all the other models (newer ones like 4o, O1, O3) seem to have knowledge cutoff October 2023? https://platform.openai.com/docs/models#o3-mini I understand that keeping the same data and curating it might be beneficial. But it sounds odd to roll back in time with the knowledge cutoff. AFAIK, the only event that happened around that time was the start of…

[flagged]
Post reply on HN