> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…
It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.
OpenAI O3-Mini
301–310 of 944 posts
Re: OpenAI O3-Mini
#302So far, it seems like this is the hierarchy o1 > GPT-4o > o3-mini > o1-mini > GPT-4o-mini o3 mini system card: https://cdn.openai.com/o3-mini-system-card.pdf
Re: OpenAI O3-Mini
#303So far, it seems like this is the hierarchy o1 > GPT-4o > o3-mini > o1-mini > GPT-4o-mini o3 mini system card: https://cdn.openai.com/o3-mini-system-card.pdf
You cannot compare GPT-4o and o*(-mini) because GPT-4o is not a reasoning model.
Re: OpenAI O3-Mini
#304Earlier quoted context omitted.
By default, we do not train on any inputs or outputs from our products for business users, including ChatGPT Team, ChatGPT Enterprise, and the API. We offer API customers a way to opt-in to share data with us, such as by providing feedback in the Playground, which we then use to improve our models. Unless they explicitly opt-in, organizations are opted out of data-sharing by default. The business bit is confusing, I…
So for posterity, in this subthread we found that OpenAI indeed trains on user data and it isn't something that only DeepSeek does.
Re: OpenAI O3-Mini
#305Earlier quoted context omitted.
So for posterity, in this subthread we found that OpenAI indeed trains on user data and it isn't something that only DeepSeek does.
So for posterity, in this subthread we found that I can use OpenAI without them training on my data, whereas I cannot with DeepSeek.
Re: OpenAI O3-Mini
#306For 18,936 input, 2,905 output it cost 3.3612 cents.
Here's the script I used to do it: https://til.simonwillison.net/llms/claude-hacker-news-themes...
Re: OpenAI O3-Mini
#307> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…
Funny - I had ChatGPT document some stuff for me this week and asked which responses I preferred as well. Didn’t bother reading either of them, just selected one and went on with my day. If it were me I would have set up a “hey do you mind if we give you two results and you can pick your favorite?” prompt to weed out people like me.
Re: OpenAI O3-Mini
#308A first step to digital immortality, could be a nice startup of some personalized product for rich, and then even regular folks. Immortality not in ourselves as meat bags of course, we die regardless, but digital copy and memento that our children can use if feeling lonely and can carry with themselves anywhere, or later descendants out of curiosity to hold massive events like weddings. One could 'invite' long lost ancestors. Maybe your grand-grand father would be a cool guy you could easily click with these days via verbal input. Heck even 3D detailed model.
An additional service, 'perpetually' paid - keeping your data model safe, taking care of it, backups, heck even maybe give it a bit of computing power to to receive current news in some light fashion and evolve, could be extras. Different tiers for different level of services and care.
Or am I decade or two ahead? I can see this as universally interesting across many if not all cultures.
Re: OpenAI O3-Mini
#309This took 1:53 in o3-mini https://chatgpt.com/share/679d310d-6064-8010-ba78-6bd5ed3360... The 4o model without using the Python tool https://chatgpt.com/share/679d32bd-9ba8-8010-8f75-2f26a792e0... Trying to get accurate results with the paid version of 4o with the Python interpreter. https://chatgpt.com/share/679d31f3-21d4-8010-9932-7ecadd0b87... The share link doesn’t show the output for some reason. But it did work…
The 4o model's output is blatantly wrong. I'm not going to look up if it's the order or the ages that are incorrect, but: 36. Abraham Lincoln – 52 years, 20 days (1861) 37. James Garfield – 49 years, 105 days (1881) 38. Lyndon B. Johnson – 55 years, 87 days (1963) Basically everything after #15 in the list is scrambled.
Re: OpenAI O3-Mini
#310> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…
It's kind of strange that they gave that stat. Maybe they thought people would somehow think about "56% better" or something. Because when you think about it, it really is quite damning. Minus statistical noise it's no better.