Live data from Hacker News

OpenAI o3-pro

help.openai.com

51–60 of 209 posts

Re: OpenAI o3-pro

#51
post #2

I understand that things are moving fast and all, but surely the.. 8? models which are currently available is a bit .. overwhelming for users that just want to get answers to their questions of life? What's the end goal with having so many models available?

I'd like one to do my test use case:

Port unix-sed from c to java with a full test suite and all options supported.

Somewhere between "it answers questions of life" and "it beats PhDs at math questions", I'd like to see one LLM take this, IMO, rather "pure" language task and succeeed.

It is complicated, but it isn't complex. It's string operations with a deep but not that deep expression system and flag set.

It is well-described and documented on the internet, and presumably training sets. It is succinctly described as a problem that virtually all computer coders would understand what it entailed if it were assigned to them. It is drudgerous, showing the opportunity for LLMs to show how they would improve true productivity.

GPT fails to do anything other than the most basic substitute operations. Claude was only slightly better, but to its detriment hallucinated massive amounts and made fake passing test cases that didn't even test the code.

The reaction I get to this test is ambivalence, but IMO if LLMs could help port entire software packages between languages with similar feature sets (aside from Turing Completeness), then software cross-use would explode, and maybe we could port "vulnerable" code to "safe" Rust en masse.

I get it, it's not what they are chasing customer-wise. They want to write (in n-gate terms) webcrap.

Re: OpenAI o3-pro

#52

I'm really hoping GPT5 is a larger jump in metrics than the last several releases we've seen like Claude3.5 - Claude4 or o3-mini-high to o3-pro. Although I will preface that with the fact I've been building agents for about a year now and despite the benchmarks only showing slight improvement, I have seen that each new generation feels actively better at exactly the same tasks I gave the previous generation. It would…

That would require AIME 2024 going above 100%. There was always going to be diminishing returns in these benchmarks. It's by construction. It's mathematically impossible for that not to happen. But it doesn't mean the models are getting better at a slower pace. Benchmark space is just a proxy for what we care about, but don't confuse it for the actual destination. If you want, you can choose to look at a different se…

There is still plenty of room for growth on the ARC-AGI benchmarks. ARC-AGI 2 is still "ARC-AGI-1: * Low: 44%, $1.64/task * Medium: 57%, $3.18/task * High: 59%, $4.16/task

ARC-AGI-2: * All reasoning efforts: Takeaways: * o3-pro in line with o3 performance * o3's new price sets the ARC-AGI-1 Frontier"

- https://x.com/arcprize/status/1932535378080395332

Re: OpenAI o3-pro

#53
post #36
post #28

The guys in the other thread who said that OpenAI might have quantized o3 and that's how they reduced the price might be right. This o3-pro might be the actual o3-preview from the beginning and the o3 might be just a quantized version. I wish someone benchmarks all of these models to check for drops in quality.

That's definitely not the case here. The new o3-pro is slow - it took two minutes just to draw me an SVG of a pelican riding a bicycle. o3-preview was much faster than that. https://simonwillison.net/2025/Jun/10/o3-pro/

Would you say this is the best cycling pelican to date? I don't remember any of the others looking better than this.

Of course by now it'll be in-distribution. Time for a new benchmark...

Re: OpenAI o3-pro

#54
post #20

Earlier quoted context omitted.

free users don't have this model selector, and probably don't care which model they get so 4o is good enough. paid users at 20$/month get more models which are better, like o3. paid users at 200$/month get the best models that are also costing OpenAI the most money, like o3-pro. I think they plan to unify them with GPT-5.

I'd be curious what proportion of paid users ever switch models. I'd guess < 10%

I switch to o1-pro on occasion, but it is slow enough that I don't use it as much as some of the others. It is a reasonably-effective last resort when I'm not getting the answer quality that I think should be achievable. It's the best available reasoning model from any provider by a noticeable margin.

Sounds like o3-pro is even slower, which is fine as long as it's better.

o4-mini-high is my usual go-to model if I need something better than the default GPT4-du jour. I don't see much point in the others and don't understand why they remain available. If o3-pro really is consistently better, it will move o1-pro into that category for me.

Re: OpenAI o3-pro

#55
post #2

I understand that things are moving fast and all, but surely the.. 8? models which are currently available is a bit .. overwhelming for users that just want to get answers to their questions of life? What's the end goal with having so many models available?

Models are used for actual tasks where predictable behavior is a benefit. Models are also used on cutting-edge tasks where smarter/better outputs are highly valued. Some applications value speed and so a new, smaller/cheaper model can be just right.

I think the naming scheme is just fine and is very straightforward to anyone who pays the slightest bit of attention.

Re: OpenAI o3-pro

#56
post #2

I understand that things are moving fast and all, but surely the.. 8? models which are currently available is a bit .. overwhelming for users that just want to get answers to their questions of life? What's the end goal with having so many models available?

I'd like one to do my test use case: Port unix-sed from c to java with a full test suite and all options supported. Somewhere between "it answers questions of life" and "it beats PhDs at math questions", I'd like to see one LLM take this, IMO, rather "pure" language task and succeeed. It is complicated, but it isn't complex. It's string operations with a deep but not that deep expression system and flag set. It is we…

How does the latest Gemini 2.5 Pro Ultra Flash Max Hemi XLT release do on that task? It obviously demands a massive context window.

Re: OpenAI o3-pro

#57

I'm really hoping GPT5 is a larger jump in metrics than the last several releases we've seen like Claude3.5 - Claude4 or o3-mini-high to o3-pro. Although I will preface that with the fact I've been building agents for about a year now and despite the benchmarks only showing slight improvement, I have seen that each new generation feels actively better at exactly the same tasks I gave the previous generation. It would…

It's hard to be 100% certain, but I am 90% certain that the benchmarks leveling off, at this point, should tell us that we are really quite dumb and simply not good very good at either using or evaluating the technology (yet?).

Re: OpenAI o3-pro

#58

[flagged]

With Gemini 2.5 in AI studio you can now increase the amount of thinking tokens, and it definitely makes a difference. O3 pro is most likely O3 with an expanded thinking token budget.

Or my favourite, tell Claude to "ultrathink"

Re: OpenAI o3-pro

#59
post #24

Earlier quoted context omitted.

At Techcrunch AI last week, the OpenAI guy started his presentation by acknowledging that OpenAI knows their naming is a problem and they're working on it, but it won't be fixed immediately.

Sam Altman has said the same thing on Twitter a few times. https://x.com/sama/status/1911906570835022319 > how about we fix our model naming by this summer and everyone gets a few more months to make fun of us (which we very much deserve) until then?

I’d prefer for them to just fix it asap instead and then keep the existing endpoints around for a year as aliases.

Re: OpenAI o3-pro

#60
post #36

Earlier quoted context omitted.

That's definitely not the case here. The new o3-pro is slow - it took two minutes just to draw me an SVG of a pelican riding a bicycle. o3-preview was much faster than that. https://simonwillison.net/2025/Jun/10/o3-pro/

Would you say this is the best cycling pelican to date? I don't remember any of the others looking better than this. Of course by now it'll be in-distribution. Time for a new benchmark...

I love that we are in the timeline where we are somewhat seriously evaluating probably super human intelligence by their ability to draw a svg of a cycling pelican.
Post reply on HN