Live data from Hacker News

OpenAI O3-Mini

openai.com

641–650 of 944 posts

Re: OpenAI O3-Mini

#641
post #530

Earlier quoted context omitted.

Yes it’s usually worth it to try to write a really good first prompt

More than once I've found myself going down this 'little maze of twisty passages, all alike'. At some point I stop, collect up the chain of prompts in the conversation, and curate them into a net new prompt that should be a bit better. Usually I make better progress - at least for a while.

This becomes second nature after a while. I've developed an intuition about when a model loses the plot and when to start a new thread. I have a base prompt I keep for the current project I'm working on, and then I ask the model to summarize what we've done in the thread and combine them to start anew.

I can't wait until this is a solved problem because it does slow me down.

Re: OpenAI O3-Mini

#642
post #3

So far, it seems like this is the hierarchy o1 > GPT-4o > o3-mini > o1-mini > GPT-4o-mini o3 mini system card: https://cdn.openai.com/o3-mini-system-card.pdf

at least if i ran the company you'd know that ChatGPTasdhjf-final-final-use_this_one.pt > ChatGPTasdhjf-final.pt > ChatGPTasdhjf.pt > ChatGPTasd.pt> ChatGPT.pt

Did you just ls my /workspace dir? Lol

Re: OpenAI O3-Mini

#643
post #80

Anyone else confused by inconsistency in performance numbers between this announcement and the concurrent system card? https://cdn.openai.com/o3-mini-system-card.pdf For example- GPQA diamond system card: o1-preview 0.68 GPQA diamond PR release: o1-preview 0.78 Also, how should we interpret the 3 different shading colors in the barplots (white, dotted, heavy dotted on top of white)...

Actually sounds like benchslop to me.

[deleted]

Re: OpenAI O3-Mini

#645
post #520

Earlier quoted context omitted.

I've found cursor to be too thin a wrapper. Aider is somehow significantly more functional. Try that.

Aider, with o1 or R1 as the architect and Claude 3.5 as the implementer, is so much better than anything you can accomplish with a single model. It's pretty amazing. Aider is at least one order of magnitude more effective for me than using the chat interface in Cursor. (I still use Cursor for quick edits and tab completions, to be clear).

Same with Cline

Re: OpenAI O3-Mini

#646
post #510

I've been using cursor since it launched, sticking almost exclusively to claude-3.5-sonnet because it is incredibly consistent, and rarely loses the plot. As subsequent models have been released, most of which claim to be better at coding, I've switched cursor to it to give them a try. o1, o1-pro, deepseek-r1, and the now o3-mini. All of these models suffer from the exact same "adhd." As an example, in a NextJS app,…

My experience with cursor and sonnet is that it is relatively good at first tries, but completely misses the plot during corrections. "My attempt at solving the problem contains a test that fails? No problem, let me mock the function I'm testing, so that, rather than actually run, it returns the expected value!" It keeps doing that kind of shenanigans, applying modifications that solve the newly appearing problem whi…

For my advanced use case involving Python and knowledge of finance, Sonnet fared poorly. Contrary to what I am reading here, my favorite approach has been to use o1 in agent mode. It’s an absolute delight to work with. It is like I’m working with a capable peer, someone at my level.

Sadly there are some hard limits on o1 with Cursor and I cannot use it anymore. I do pay for their $20/month subscription.

Re: OpenAI O3-Mini

#647
post #510

I've been using cursor since it launched, sticking almost exclusively to claude-3.5-sonnet because it is incredibly consistent, and rarely loses the plot. As subsequent models have been released, most of which claim to be better at coding, I've switched cursor to it to give them a try. o1, o1-pro, deepseek-r1, and the now o3-mini. All of these models suffer from the exact same "adhd." As an example, in a NextJS app,…

So I think I finally understood recently why we have these divergent groups with one thinking Claude 3.5 Sonnet is the best model for coding and another that follow the OpenAI SOTA at that moment. I have been a heavy user of ChatGPT, jumping on to pro without even thinking for more than a second once released. Recently though I took a pause from my usual work on statistical modelling, heuristics work and other things in certain deep domains to focus on building client APIs and frontends and decided to again give Claude a try and it is just so great to work with for this usecase.

My hypothesis is its a difference of what you are doing. OpenAI O models are much better than others at mathematical modelling and such tasks and Claude for more general purpose programming.

Re: OpenAI O3-Mini

#648

Earlier quoted context omitted.

When asking it about the concept of human rights and the various forms in which it manifests (i.e. demographic equality under the law). I get a mixture of mundane nuance and bizarre answers that Xi Jingping himself could have written. With references to unity and the importance of social harmony over the "freedoms of the few". This tracks when considering that the model was trained on western model outputs and then t…

I definitely am not getting that, perhaps the 671b model is notably worse than the 70b llama distill in this respect. 70b seemed pretty happy to talk about the ethnic cleansing of the Uyghurs in Xinjiang by the CCP and Palestinians in Gaza by Israel, it did some both-sides ing but it generally seemed to provide a balanced-ish viewpoint. At least I think it provided a viewpoint that comports with my best guess of what…

My favorite experience with the 70b distill was to ask it why communism consistently resulted in mass murder. It gave an immediate boilerplate response saying it doesn't and glorifying the Chinese communist party, then went into think mode and talked itself into the position that communism has, in fact, consistently resulted in resulted in mass murder.

They have under utilized the chain of thought in their resoning, it ought to be thinking something like "I need to be careful to not say anything that could bring embarrassment to the party"..

but perhaps the online versions do actually preload the reasoning this way. :P

Re: OpenAI O3-Mini

#649

> Testers preferred o3-mini's responses to o1-mini 56% of the time I hope by this they don't mean me, when I'm asked 'which of these two responses do you prefer'. They're both 2,000 words, and I asked a question because I have something to do. I'm not reading them both ; I'm usually just selecting the one that answered first. That prompt is pointless. Perhaps as evidenced by the essentially 50% response rate: it's a…

> I'm usually just selecting the one that answered first Which is why you randomize the order. You aren’t a tester. 56% vs 44% may not be noise. That’s why we have p values. It depends on sample size.

The order doesn't matter. They often generate tokens at different speeds, and produce different lengths of text. "The one that answered first" != "The first option"

Re: OpenAI O3-Mini

#650

Earlier quoted context omitted.

A 12% margin is literally the opposite of a coin flip. Unless you have a really bad coin.

You're being downvoted for 3 reasons: 1) Coming off as a jerk, and from a new account is a bad look 2) "Literally the opposite of a coin flip" would probably be either 0% or 100% 3) Your reasoning doesn't stand up without further info; it entirely depends on the sample size. I could have 5 coin flips all come up heads, but over thousands or millions it averages to 50%. 56% on a small sample size is absolutely within…

I'm a little puzzled by your response.

1. The message was net-upvoted. Whether there are downvotes in there I can't tell, but the final karma is positive. A similarly spirited message of mine in the same thread was quite well receive as well.

2. I can't see how my message would come across as a jerk? I wrote 2 simple sentences, not using any offensive language, stating a mere fact of statistics. Is that being jerk? And a long-winded berating of a new member of the community isn't?

3. A coin flip is 50%. Anything else is not, once you have a certain sample size. So, this was not. That was my statement. I don't know why you are building a strawman of 5 coin flips. 56% vs 44% is a margin of 12%, as I stated, and with a huge sample size, which they had, that's massive in a space where the returns are deep in "diminishing" territory.

Post reply on HN