Live data from Hacker News

OpenAI o3-pro

help.openai.com

41–50 of 209 posts

Re: OpenAI o3-pro

#41
post #39
post #36

Earlier quoted context omitted.

That's definitely not the case here. The new o3-pro is slow - it took two minutes just to draw me an SVG of a pelican riding a bicycle. o3-preview was much faster than that. https://simonwillison.net/2025/Jun/10/o3-pro/

Not distilled, same model. https://x.com/therealadamg/status/1932534244774957121?s=46&t...

[deleted]

Re: OpenAI o3-pro

#42
post #28

The guys in the other thread who said that OpenAI might have quantized o3 and that's how they reduced the price might be right. This o3-pro might be the actual o3-preview from the beginning and the o3 might be just a quantized version. I wish someone benchmarks all of these models to check for drops in quality.

o3-pro is not the same as the o3-preview that was shown in Dec '24. OpenAI confirmed this for us. More on that here: https://x.com/arcprize/status/1932535380865347585

Re: OpenAI o3-pro

#43

I'm really hoping GPT5 is a larger jump in metrics than the last several releases we've seen like Claude3.5 - Claude4 or o3-mini-high to o3-pro. Although I will preface that with the fact I've been building agents for about a year now and despite the benchmarks only showing slight improvement, I have seen that each new generation feels actively better at exactly the same tasks I gave the previous generation. It would…

I'm seeing big advances that arent shown in the benchmarks, I can simply build software now that I couldnt build before. The level of complexity that I can manage and deliver is higher.

Yeah I kind of feel like I'm not moving as fast as I did, because the complexity and features grow - constant scope creep due to moving faster.

Re: OpenAI o3-pro

#44

I'm really hoping GPT5 is a larger jump in metrics than the last several releases we've seen like Claude3.5 - Claude4 or o3-mini-high to o3-pro. Although I will preface that with the fact I've been building agents for about a year now and despite the benchmarks only showing slight improvement, I have seen that each new generation feels actively better at exactly the same tasks I gave the previous generation. It would…

That would require AIME 2024 going above 100%.

There was always going to be diminishing returns in these benchmarks. It's by construction. It's mathematically impossible for that not to happen. But it doesn't mean the models are getting better at a slower pace.

Benchmark space is just a proxy for what we care about, but don't confuse it for the actual destination.

If you want, you can choose to look at a different set of benchmarks like ARC-AGI-2 or Epoch and observe greater than linear improvements, and forget that these easier benchmarks exist.

Re: OpenAI o3-pro

#45
post #28

The guys in the other thread who said that OpenAI might have quantized o3 and that's how they reduced the price might be right. This o3-pro might be the actual o3-preview from the beginning and the o3 might be just a quantized version. I wish someone benchmarks all of these models to check for drops in quality.

Is there a way to figure out likely quantization from the output. I mean, does quantization degrade output quality in certain ways that are different from other modification of other model properties (e.g. size or distillation)?

Re: OpenAI o3-pro

#46
post #35

Earlier quoted context omitted.

It is not thinking. It is trying to deceive you. The ”reasoning” it outputs does not have a causal relationship with the end result.

The longer "it" reasons, the more attention sinks are used to come to a "better" final output.

I’ve looked up attention sinks and can’t figure out how you’re using the term here. It sounds interesting, would you care to elaborate?

Re: OpenAI o3-pro

#47

Earlier quoted context omitted.

it's the same model as o3, just with thinking tokens turned up to the max.

That's simply not true, it's not just "max thinking budget o3" just like o1-pro wasn't "max thinking budget o1". The specifics are unknown, but they might be doing multiple model generations and then somehow picking the best answer each time? Of course that's a gross simplification, but some assume that they do it this way.

> "We also introduced OpenAI o3-pro in the API—a version of o3 that uses more compute to think harder and provide reliable answers to challenging problems"

Sounds like it is just o3 with higher thinking budget to me

Re: OpenAI o3-pro

#48

Earlier quoted context omitted.

I just can’t believe nobody at the company has enough courage to tell their leadership that their naming scheme is completely stupid and insane. Four is greater than three, and so four should be better than three. The point of a name is to describe something so that you don’t confuse your users, not to be cute.

What’s worse is that the app doesn’t even have descriptions. As if I’m supposed to memorize the use case for each based on: GPT-4o o3 o4-mini o4-mini-high GPT-4.5 GPT-4.1 GPT-4.1-mini

Just use o4-mini for everything

Re: OpenAI o3-pro

#49

Earlier quoted context omitted.

> That's simply not true, it's not just "max thinking budget o3" > The specifics are unknown, but they might... Hold up. > but some assume that they do it this way. Come on now.

Good luck finding the tweet (I can't) but at least one OpenAI engineer has said that o1-pro was not just 'o1 thinking longer'.

This one? Found with Kagi Assistant.

https://x.com/michpokrass/status/1869102222598152627

It says:

> hey aidan, not a miscommunication, they are different products! o1 pro is a different implementation and not just o1 with high reasoning.

Re: OpenAI o3-pro

#50
post #13

here's a nice user review we published: https://www.latent.space/p/o3-pro sama's highlight[0]: > "The plan o3 gave us was plausible, reasonable; but the plan o3 Pro gave us was specific and rooted enough that it actually changed how we are thinking about our future." I kept nudging the team to go the whole way to just let o3 be their CEO but they didn't bite yet haha 0: https://x.com/sama/status/1932533208366608568

Big fan swyx, but both here and in the article there is some bragging about being quoted by sama, and while I acknowledge that that’s not out of the ordinary, I’m concerned about where it leads: what it takes to get quoted by sama (or similar interested party) is saying something good about his product, and having a decent follower count.

Dangerous incentives IMO.

Post reply on HN