Live data from Hacker News

OpenAI O3-Mini

openai.com

941–944 of 944 posts

Re: OpenAI O3-Mini

#941

Earlier quoted context omitted.

Thanks for the link. >> It effectively treats “reasoning” as the ability to generate intermediate steps leading to a correct conclusion. Is "effectively" the same as "pretty precise" as per your previous comment? I don't see that because I searched the paper for all occurrences of "reasoning" and noticed two things: first that while the term is used to saturation there is no attempt to define it even informally, let…

The point is there has to be meaning for reasoning. I think the claim in this paper is very clear and the results are shown decisively. Research papers relating to reasoning approach and define it in many ways but crucially, the good ones offer a testable claim. Simply saying “models can’t reason” is ambiguous to the point of being unanswerable.

[deleted]

Re: OpenAI O3-Mini

#942

Earlier quoted context omitted.

You're conflating the low price of the o3-mini medium effort model with the high performance of the o3-mini high effort model. OpenAI hasn't listed the price for the o3-mini high effort model separately on their pricing page.

If they are the same underlying model, it’s unlikely the prices will be different on a per token basis. The high model will simply consume more tokens.

You're right but then in that mode it's no longer cheap.

Re: OpenAI O3-Mini

#943

Earlier quoted context omitted.

Also Gemini API is free for coding.

I have yet to see a valid reason to use Gemini over any other alternative, the only exception being large contexts.

2 million tokens context windows separates it from all other LLMs.

Re: OpenAI O3-Mini

#944

Earlier quoted context omitted.

More than once I've found myself going down this 'little maze of twisty passages, all alike'. At some point I stop, collect up the chain of prompts in the conversation, and curate them into a net new prompt that should be a bit better. Usually I make better progress - at least for a while.

This becomes second nature after a while. I've developed an intuition about when a model loses the plot and when to start a new thread. I have a base prompt I keep for the current project I'm working on, and then I ask the model to summarize what we've done in the thread and combine them to start anew. I can't wait until this is a solved problem because it does slow me down.

Yes when new models come out it feels like breaking up.
Post reply on HN