Live data from Hacker News

OpenAI O3-Mini

openai.com

761–770 of 944 posts

Re: OpenAI O3-Mini

#761
post #626

Earlier quoted context omitted.

Why is it so hard to share/find prompts or distill my own damn prompts? There must be good solutions for this —

What do you find difficult about distilling your own prompts? After any back and forth session I have reasonably good results asking something like "Given this workflow, how could I have prompted this better from the start to get the same results?"

Analysis of past chats in bulk.

Re: OpenAI O3-Mini

#762
post #237

Earlier quoted context omitted.

Maybe they found a need to quantize it further for release, or lobotomise it with more "alignment".

> lobotomise Anyone can write very fast software if you don't mind it sometimes crashing or having weird bugs. Why do people try to meme as if AI is different? It has unexpected outputs sometimes, getting it to not do that is 50% "more alignment" and 50% "hallucinate less". Just today I saw someone get the Amazon bot to roleplay furry erotica. Funny, sure, but it's still obviously a bug that a *sales bot* would do th…

Who determines who gets access to what information? The OpenAI board? Sam? What qualifies as dangerous information? Maybe it’s dangerous to allow the model to answer questions about a person. What happens when limiting information becomes a service you can sell? For the right price anything can become too dangerous for the average person to know about.

Re: OpenAI O3-Mini

#763

Earlier quoted context omitted.

The point is that you cannot make a general statement that “o1 is better than 4o.”

Yes, but because you need to say exactly what one is better than the other for. Not because o1 spends a bunch of tokens for "reasoning" you cannot even see.

If you would like to see the CoT process visualized, try the “Improve prompt” feature in Anthropic console. Also check out https://github.com/getAsterisk/deepclaude

Re: OpenAI O3-Mini

#765
post #669
post #306

I used o3-mini to summarize this thread so far. Here's the result: https://gist.github.com/simonw/09e5922be0cbb85894cf05e6d75ae... For 18,936 input, 2,905 output it cost 3.3612 cents. Here's the script I used to do it: https://til.simonwillison.net/llms/claude-hacker-news-themes...

Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.

Amazon is already forcing this pattern on mobile users not logged in. If you want to see the reviews, all you get is an AI summary, the star rating, maybe 1 or 2 reviews, then you have to log in to see more.

Re: OpenAI O3-Mini

#766
post #669

Earlier quoted context omitted.

Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.

But then there will be no comments to summarize.

But what if the AI just hallucinates the comments? People will never know.

Re: OpenAI O3-Mini

#767
post #669

Earlier quoted context omitted.

Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.

Amazon is already forcing this pattern on mobile users not logged in. If you want to see the reviews, all you get is an AI summary, the star rating, maybe 1 or 2 reviews, then you have to log in to see more.

The cure is to stop using Amazon.

Re: OpenAI O3-Mini

#768
i think it says, amongst other things, that there is a salient difference between competitive programming like codeforce and real-world programming. u can train a model to hillclimb elo ratings on codeforce, but that won't necessarily directly translate to working on a prod javascript codebase.

anthropic figured out something about real world coding that openai is still trying to catch up to, o3-mini-high notwithstanding.

Re: OpenAI O3-Mini

#770

After o3 was announced, with the numbers suggesting it was a major breakthrough, I have to say I’m absolutely not impressed with this version. I think o1 works significantly better, and that makes me think the timing is more than just a coincidence. Last week Nvidia lost 600 billion because of DeepSeek R1, and now OpenAI comes out with a new release which feels like it has nothing to do with the promises that were be…

This is the mini version which is not as good as o1 and I don’t think they demoed in the o3 announcement. I’m hoping the full release will be impressive
Post reply on HN