When I use ChatGPT these days, it's to help me write coding videos and then the social media posts around those videos. So that's two specialties in one chat log.
OpenAI O3-Mini
491–500 of 944 posts
Re: OpenAI O3-Mini
#492Earlier quoted context omitted.
It's definitely making some errors
Like?
> It’s already a game changer for many people. But to have so many names like o1, o3-mini, GPT-4o, & GPT-4o-mini suggests there may be too much focus on internal tech details rather than clear communication." (paraphrase based on multiple similar sentiments)
It also hallucinates quotes.
For example:
> "I’m pretty sure 'o3-mini' works better for that purpose than 'GPT 4.1.3'." – TeMPOraL
But that comment is not in the user TeMPOraL's comment history.
Sentiment analysis is also faulty.
For example:
> "I’d bet most users just 50/50 it, which actually makes it more remarkable that there was a 56% selection rate." – jackbrookes – This quip injects humor into an otherwise technical discussion about evaluation metrics.
It's not a quip though. That comment was meant in earnest
Re: OpenAI O3-Mini
#493Earlier quoted context omitted.
Funny - I had ChatGPT document some stuff for me this week and asked which responses I preferred as well. Didn’t bother reading either of them, just selected one and went on with my day. If it were me I would have set up a “hey do you mind if we give you two results and you can pick your favorite?” prompt to weed out people like me.
I'm surprised how many people claim to do this. You can just not select one.
We find what appears to be the easiest combination "Fuck off, go away" buttons and use them without a moment of actual consideration.
(This doesn't mean that it's actually the easiest method.)
Re: OpenAI O3-Mini
#494Earlier quoted context omitted.
I like deepseek a lot. But they are currently very glitchy. The API service goes up and down a lot. Maybe they'll sort that out soon.
Apparently they're under a very targeted DDoS for almost a month, with technical details shared in Chinese but very little discussion in English. Which is surprising, it's not like major AI products are getting DDoSed out of existence every day.
Re: OpenAI O3-Mini
#495Earlier quoted context omitted.
I don't want a model that's customized to my preferences. My preferences and understanding changes all the time. I want a single source model that's grounded in base truth. I'll let the model know how to structure it in my prompt.
You know there's no such as base truth here? You want to write something like this to start your prompts, "Respond in English, using standard capitalization and punctuation, following rules of grammar as written by Strunk & White, where numbers are represented using arabic numerals in base 10 notation...."???
I like precision of language, so maybe just have a system prompt that says "use precise language (ex: no symbolism of any kind)"
Re: OpenAI O3-Mini
#496O3-mini solved this prompt. DeepSeek R1 had a mental breakdown. The prompt: “Bob is facing forward. To his left is Ann, to his right is Cathy. Ann and Cathy are facing backwards. Who is on Ann’s left?”
Re: OpenAI O3-Mini
#497Earlier quoted context omitted.
Typically in these tests you have three options "A is better", "B is better" or "they're equal/can't decide". So if 56% prefer O3 Mini, it's likely that way less than half prefer O1.also, the way I understand it, they're comparing a mini model with a large one.
If you use ChatGPT, it sometimes gives you two versions of its response, and you have to choose one or the other if you want to continue prompting. Sure, not picking a response might be a third category. But if that's how they were approaching the analysis, they could have put out a more favorable-looking stat.
Does no one else hate it when this happens (especially when on a handheld device)?
Re: OpenAI O3-Mini
#498Can't wait to try this. What's amazing to me is that when this was revealed just one short month ago, the AI landscape looked very different than it does today with more AI companies jumping into the fray with very compelling models. I wonder how the AI shift has affected this release internally, future releases and their mindset moving forward... How does the efficiency change, the scope of their models, etc.
I thought it was o3 that was released one month ago and received high scores on ARC Prize - https://arcprize.org/blog/oai-o3-pub-breakthrough If they were the same, I would have expected explicit references to o3 in the system card and how o3-mini is distilled or built from o3 - https://cdn.openai.com/o3-mini-system-card.pdf - but there are no references. Excited at the pace all the same. Excited to dig in. The model…
Re: OpenAI O3-Mini
#499> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…
It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.
Re: OpenAI O3-Mini
#500Earlier quoted context omitted.
How would the DeepSeek fit into this? Or can it not compare? I don't know much about this stuff, but I've heard recently many people talk about DeepSeek and how unexpected it was.
Deepseek V3 is equivalent to 4o. Deepseek R1 is equivalent to o1 (if not better) I think someone should just build an AI model comparing website at this point. Include all benchmarks and pricing
One thing I notice that is huge is being able to see the chain of thought lets me see when my prompt was lacking and the model is a bit confused on what I want.
If I was anymore impressed with R1 I would probably start getting accused of being a CCP shill or wumao lol.
With that said, I think it is very hard to compare models for your own use case. I do suspect there is a shiny new toy bias with all this too.
Poor Sonnet 3.5. I have neglected it so much lately I actually don't know if I have a subscription or not right now.
I do expect an Anthropic reasoning model though to blow everything else away.