Live data from Hacker News

OpenAI O3-Mini

openai.com

241–250 of 944 posts

Re: OpenAI O3-Mini

#241

What is the comparison of this versus DeepSeek in terms of good results and cost?

Deepseek is the state of the art right now in terms of performance and output. It's really fast. The way it "explains" how it's thinking is remarkable.

DeepSeek is great because: 1) you can run the model locally, 2) the research was openly shared, and 3) the reasoning tokens are open. It is not, in my experience, state of the art. In all of my side by side comparisons thus far in real world applications between DeepSeek V3 and R1 vs 4o and o1, the latter has always performed better. OpenAI's models are also more consistent, glitching out maybe one in 10,000, whereas DeepSeek's models will glitch out 1 in 20. OpenAI models also handle edge cases better and have a better overall grasp of user intentions. I've had DeepSeek's models consistently misinterpret prompts, or confuse data in the prompts with instructions. Those are both very important things that make DeepSeek useless for real world applications. At least without finetuning them, which then requires using those huge 600B parameter models locally.

So it is by no means state of the art. Gemini Flash 2.0 also performs better than DeepSeek V3 in all my comparisons thus far. But Gemini Flash 2.0 isn't robust and reliable either.

But as a piece of research, and a cool toy to play with, I think DeepSeek is great.

Re: OpenAI O3-Mini

#243
post #25

I think OpenAI should just have a single public facing "model" - all these names and versions are confusing. Imagine if Google, during it's accent, had a huge array of search engines with code names and notes about what it's doing behind the scenes. No, you open the page and type in box. If they can make it work better next month, great. (I understand this could not apply to developers or enterprise-type API usage).

If google had to face the reality that distilling their search engine into multiple case-specific engines would have resulted in vastly superior search results, they surely would done (or considered) it.

Fortunately for them a monolith search engine was perfectly fine (and likely optimal due to accrued network effects).

OpenAI is basically signaling that they need to distill their monolith in order to serve specific segments of the marketplace. They've explicitly said that they're targeting STEM with this one. I think that's a smart choice, the most passionate early adopters of this tech are clearly STEM users.

If the tech was such that one monolith model was actually the optimal solution for all use cases, they would just do that. Actually, this is their stated mission: AGI. One monolith that's best at everything is basically what AGI is.

Re: OpenAI O3-Mini

#244
post #196

I think that OpenAI should reduce the prices even further to be competitive with Qwen or Deepseek. There are a lot of vendors offering Deepseek R1 for $2-2.5 per 1 million tokens output.

Would you have specific recommendations of such vendors?

For example, `https://deepinfra.com/` which asks for $2.5 per million on output or https://nebius.com which asks for $2.4 per million output tokens.

Re: OpenAI O3-Mini

#246

The naming convention is so messed up. o1, o3-mini (no o2, no o3???)

https://www.perplexity.ai/search/new?q=list%20of%20all%20Ope... :)

OpenAI has developed a variety of models that cater to different applications, from natural language processing to image generation and audio processing. Here’s a comprehensive list of the current models available:

   ## Language Models
   - \*GPT-4o\*: The flagship model capable of processing text, images, and audio.
   - \*GPT-4o mini\*: A smaller, more cost-effective version of GPT-4o.
   - \*GPT-4\*: An advanced model that improves upon GPT-3.5.
   - \*GPT-3.5\*: A set of models that enhance the capabilities of GPT-3.
   - \*GPT-3.5 Turbo\*: A faster variant designed for efficiency in chat applications.

   ## Reasoning Models
   - \*o1\*: Focused on reasoning tasks with improved accuracy.
   - \*o1-mini\*: A lightweight version of the o1 model.
   - \*o3\*: The successor to o1, currently in testing phases.
   - \*o3-mini\*: A lighter version of the o3 model.

   ## Audio Models
   - \*GPT-4o audio\*: Supports real-time audio interactions and audio generation.
   - \*Whisper\*: For transcribing and translating speech to text.

   ## Image Models
   - \*DALL-E\*: Generates images from textual descriptions.

   ## Embedding Models
   - \*Embeddings\*: Converts text into numerical vectors for similarity tasks.
   - \*Ada\*: An embedding model with various sizes (e.g., ada-002).

   ## Additional Models
   - \*Text to Speech (Preview)\*: Synthesizes spoken audio from text.
These models are designed for various tasks, including coding assistance, image generation, and conversational AI, making OpenAI's offerings versatile for developers and businesses alike[1][2][4][5].

Citations:

   [1] https://learn.microsoft.com/vi-vn/azure/ai-services/openai/concepts/models
   [2] https://platform.openai.com/docs/models
   [3] https://llm.datasette.io/en/stable/openai-models.html
   [4] https://en.wikipedia.org/wiki/OpenAI_API
   [5] https://industrywired.com/open-ai-models-list-top-models-to-consider/
   [6] https://holypython.com/python-api-tutorial/listing-all-available-openai-models-openai-api/
   [7] https://en.wikipedia.org/wiki/GPT-3
   [8] https://stackoverflow.com/questions/78122648/openai-api-how-do-i-get-a-list-of-all-available-openai-models/78122662

Re: OpenAI O3-Mini

#247

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

RLUHF, U = useless.

Re: OpenAI O3-Mini

#248
post #196

I think that OpenAI should reduce the prices even further to be competitive with Qwen or Deepseek. There are a lot of vendors offering Deepseek R1 for $2-2.5 per 1 million tokens output.

Would you have specific recommendations of such vendors?

Well, it's $2.19 per million output tokens even directly on deepseek platform.

https://api-docs.deepseek.com/quick_start/pricing/

Re: OpenAI O3-Mini

#249

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

Those prompts are so irritating and so frequent that I’ve taken to just quickly picking whichever one looks worse at a cursory glance. I’m paying them, they shouldn’t expect high quality work from me.
Post reply on HN