Live data from Hacker News

OpenAI o3-pro

help.openai.com

181–190 of 209 posts

Re: OpenAI o3-pro

#181
post #67

Earlier quoted context omitted.

I remember the saying that from 90% to 99% is a 10x increase in accuracy, but 99% to 99.999% is a 1000x increase in accuracy. Even though it's a large10% increase first then only a 0.999% increase.

Sometimes it’s nice to frame it the other way, eg: 90% -> 1 error per 10 99% -> 1 error per 100 99.99% -> 1 error per 10,000 That can help to see the growth in accuracy, when the numbers start getting small (and why clocks are framed as 1 second lost per…).

Still, for the human mind it doesn't make intuitive sense.

I guess it's the same problem with the mind not intuitively grasping the concept of exponential growth and how fast it grows.

Re: OpenAI o3-pro

#182
post #181

Earlier quoted context omitted.

Sometimes it’s nice to frame it the other way, eg: 90% -> 1 error per 10 99% -> 1 error per 100 99.99% -> 1 error per 10,000 That can help to see the growth in accuracy, when the numbers start getting small (and why clocks are framed as 1 second lost per…).

Still, for the human mind it doesn't make intuitive sense. I guess it's the same problem with the mind not intuitively grasping the concept of exponential growth and how fast it grows.

ChatGPT quick explanation:

Humans struggle with understanding exponential growth due to a cognitive bias known as *Exponential Growth Bias (EGB)*—the tendency to underestimate how quickly quantities grow over time. Studies like Wagenaar & Timmers (1979) and Stango & Zinman (2009) show that even educated individuals often misjudge scenarios involving doubling, such as compound interest or viral spread. This is because our brains are wired to think linearly, not exponentially, a mismatch rooted in evolutionary pressures where linear approximations were sufficient for survival.

Further research by Tversky & Kahneman (1974) explains that people rely on mental shortcuts (heuristics) when dealing with complex concepts. These heuristics simplify thinking but often lead to systematic errors, especially with probabilistic or nonlinear processes. As a result, exponential trends—such as pandemics, technological growth, or financial compounding—often catch people by surprise, even when the math is straightforward.

Re: OpenAI o3-pro

#183
post #102

Earlier quoted context omitted.

"most people I show them too have issues understanding them, and in fact I had issues understanding them" ??? those benchmarks are so extremely simple they have basically 100% human approval rates, unless you are saying "I could not grasp it immediately but later I was able to after understanding the point" I think you and your friends should see a neurologist. And I'm not mocking you, I mean seriously, those are tas…

> so extremely simple they have basically 100% human approval rates Are you thinking of a different set? Arc-agi-2 has average 60% success for a single person and questions require only 2 out of 9 correct answers to be accepted. https://docs.google.com/presentation/d/1hQrGh5YI6MK3PalQYSQs... > and even some other mammals to do. No, that's not the case.

No, I think I saw the graphs on someone's channel, but maybe I misinterpreted the results. But to be fair, my point never depended on 100% of the participants being right 100% of the questions, there are innumerous factors that could affect your performance on those tests, including the pressure. The AI also had access to lenient conventions, so it should be "fair" in this sense.

Either way, there's something fishy about this presentation, it says: "ARC-AGI-1 WAS EASILY BRUTE-FORCIBLE", but when o3 initially "solved" most of it the co-founder or ARC-PRIZE said: "Despite the significant cost per task, these numbers aren't just the result of applying brute force compute to the benchmark. OpenAI's new o3 model represents a significant leap forward in AI's ability to adapt to novel tasks. This is not merely incremental improvement, but a genuine breakthrough, marking a qualitative shift in AI capabilities compared to the prior limitations of LLMs. o3 is a system capable of adapting to tasks it has never encountered before, arguably approaching human-level performance in the ARC-AGI domain.", he was saying confidently that it would not be a result of brute-forcing the problems. And it was not the first time, "ARC-AGI-1 consists of 800 puzzle-like tasks, designed as grid-based visual reasoning problems. These tasks, trivial for humans but challenging for machines, typically provide only a small number of example input-output pairs (usually around three). This requires the test taker (human or AI) to deduce underlying rules through abstraction, inference, and prior knowledge rather than brute-force or extensive training."

Now they are saying ARC-AGI-2 is not bruteforcible, what is happening there? They didn't provided any reasoning for why one was bruteforcible and the other not, nor how they are so sure about that. They "recognized" that it could be brute-forced before, but in a way less expressive manner, by explicitly stating it would need "unlimited resources and time" to solve. And they are using the non-bruteforceability in this presentation as a point for it.

--- Also, I mentioned mammals because those problems are of an order that mammals and even other animals would need to solve in reality for a diversity of cases. I'm not saying that they would literally be able to take the test and solve it, nor to understand this is a test, but that they would need to solve problems of similar nature in reality. Naturally this point has it's own limits, but it's not easily discarded as you tried to do.

Re: OpenAI o3-pro

#184
post #102

Earlier quoted context omitted.

"most people I show them too have issues understanding them, and in fact I had issues understanding them" ??? those benchmarks are so extremely simple they have basically 100% human approval rates, unless you are saying "I could not grasp it immediately but later I was able to after understanding the point" I think you and your friends should see a neurologist. And I'm not mocking you, I mean seriously, those are tas…

lol 100% approval rates? No they don’t. Also mammals? What mammals could even understand we were giving it a test? Have you seen them or shown them to average people? I’m sure the people who write them understand them but if you show these problems to average people in the street they are completely clueless. This is a classic case of some phd ai guys making a benchmark and not really considering what average people…

quoting my own previous response: > Also, I mentioned mammals because those problems are of an order that mammals and even other animals would need to solve in reality for a diversity of cases. I'm not saying that they would literally be able to take the test and solve it, nor to understand this is a test, but that they would need to solve problems of similar nature in reality. Naturally this point has it's own limits, but it's not easily discarded as you tried to do.

---

> Have you seen them or shown them to average people? I’m sure the people who write them understand them but if you show these problems to average people in the street they are completely clueless.

I can show them to people on my family, I'll do it today and come back with the answer, it's the best way of testing that out.

Re: OpenAI o3-pro

#185

Earlier quoted context omitted.

Can you provide the details? Sounds intriguing

The frustrating thing about private problems like this is that if you respond to requests like this, it'll become part of the training data. I'm fairly certain HN is scraped because several AIs know my HN alias and can replicate my style of writing on demand. PS: Thinking about it... that is a very specific kind of disturbing feeling that only prolific online commenters can experience... There's a soulless machine so…

Perhaps we are not as unique as we’d like to believe.

The machine does not understand you. The machine can match your flavor of textual communication.

This can be done for audio with a relatively small number of samples. Your iPhone has a feature called Personal Voice which claims it can do it with 150 phrases/15 minutes of your time.

Re: OpenAI o3-pro

#186
post #183

Earlier quoted context omitted.

> so extremely simple they have basically 100% human approval rates Are you thinking of a different set? Arc-agi-2 has average 60% success for a single person and questions require only 2 out of 9 correct answers to be accepted. https://docs.google.com/presentation/d/1hQrGh5YI6MK3PalQYSQs... > and even some other mammals to do. No, that's not the case.

No, I think I saw the graphs on someone's channel, but maybe I misinterpreted the results. But to be fair, my point never depended on 100% of the participants being right 100% of the questions, there are innumerous factors that could affect your performance on those tests, including the pressure. The AI also had access to lenient conventions, so it should be "fair" in this sense. Either way, there's something fishy a…

> my point never depended on 100% of the participants being right 100% of the questions

You told someone that their reasoning is so bad they should get checked by a doctor. Because they didn't find the test easy, even though it averages 60% score per person. You've been a dick to them while significantly misrepresenting the numbers - just stop digging.

Re: OpenAI o3-pro

#187
post #94

Earlier quoted context omitted.

I've been using o3 extensively since release (and a lot of Deep Research). I also use a lot of Claude and Gemini 2.5 Pro (most of the times, for code I'll let all of them go at it and iterate on my fav results). So far I've only used o3-pro a bit today, and it's a bit too heavy to use interactively (fire it off, revisit in 10-15 minutes), but it seems to generate much cleaner/more well organized code and answers. I f…

I wonder if we'll start to see artisanal benchmarks. You -- and I -- have preferred models for certain tasks. There's a world in which we start to see how things score on the "simonw chattiness index", and come to rely on smaller more specific benchmarks I think

I think its more likely that we move away from benchmarks and towards more of a traditional reviewer model. People will find LLM influencers whose takes they agree with and follow them to keep up with new models.

Re: OpenAI o3-pro

#188
post #183

Earlier quoted context omitted.

No, I think I saw the graphs on someone's channel, but maybe I misinterpreted the results. But to be fair, my point never depended on 100% of the participants being right 100% of the questions, there are innumerous factors that could affect your performance on those tests, including the pressure. The AI also had access to lenient conventions, so it should be "fair" in this sense. Either way, there's something fishy a…

> my point never depended on 100% of the participants being right 100% of the questions You told someone that their reasoning is so bad they should get checked by a doctor. Because they didn't find the test easy, even though it averages 60% score per person. You've been a dick to them while significantly misrepresenting the numbers - just stop digging.

The second test scores 60%, the first was way higher. And I specifically said ""unless you are saying "I could not grasp it immediately but later I was able to after understanding the point" I think you and your friends should see a neurologist"", to which this person did not responded. I saw the tests, solved some, I suspect the variability here is more a question of methodology than an inherent problem for those people. I also never stated that my point depended on those people scoring 100% specifically on the tests, even if it is in fact extremely easy (and it is, the objective of this test is to literally make tests that most humans could easily beat but that would be hard for an AI) variability will still exist and people with different perceptions would skew the results, this is expected. "Significantly misrepresenting the numbers" is also a stretch, I only mentioned the numbers ONE time in my point, most of it was about that inherent nature (or at least, the intended nature) of the tests.

So on the edge, if he was not able to understand them at all, and this was not just a problem of grasping the problem, my point was that this would possibly indicate a neurological problem, or developmental, due to the nature of them. It's not a question of "you need to get all of them right", his point was that he was unable to understand them at all, that it confused them to an understanding level.

Re: OpenAI o3-pro

#189
post #20

Earlier quoted context omitted.

I'd be curious what proportion of paid users ever switch models. I'd guess < 10%

If you're not at least switching from 4o to 4.1 you're doing it wrong.

4o is better than 4.1 for a lot of things that are non-coding/general research.

Re: OpenAI o3-pro

#190
post #181

Earlier quoted context omitted.

Sometimes it’s nice to frame it the other way, eg: 90% -> 1 error per 10 99% -> 1 error per 100 99.99% -> 1 error per 10,000 That can help to see the growth in accuracy, when the numbers start getting small (and why clocks are framed as 1 second lost per…).

Still, for the human mind it doesn't make intuitive sense. I guess it's the same problem with the mind not intuitively grasping the concept of exponential growth and how fast it grows.

The lily pad example of the lake being half full on the 29th day out of 30 is also a good one.
Post reply on HN