Live data from Hacker News

Epoch confirms GPT5.4 Pro solved a frontier math open problem

epoch.ai

211–220 of 744 posts

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#211

Earlier quoted context omitted.

I doubt you can even define intelligence sufficiently to argue this point. Since that's an ongoing debate without a resolution thus far. But you claimed that humans aren't unique. I think it's pretty obvious we are on many dimensions including what you might classify as "intelligence". You don't even necessarily have to believe in a "soul" or something like that, although many people do. The capabilities of a human f…

> I often wonder why tech has so many reductionist, materialist, and quite frankly anti-human, thinkers. I think it comes from a position of arrogance/ego. I'll speak for the US here, since that's what I know the most; but the average 'techie' in general skews towards the higher intelligence numbers than the lower parts. This is a very, very broad stroke, and that's intentional to illustrate my point. Because of this…

I largely agree with you, but I also see this same type of thinking appear in people who I know are not arrogant - at least in the techbroisk way.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#212

I have long said I am an AI doubter until AI could print out the answers to hard problems or ones requiring tons of innovation. Assuming this is verified to be correct (not by AI) then I just became a believer. I would like to see a few more AI inventions to know for sure, but wow, it really is a new and exciting world. I really hope we use this intelligence resource to make the world better.

Are the only two options AI doubter and AI believer?

Perhaps I should have elaborated more but what I mean is I used to think, "I genuinely don't see the point in even trying to use AI for things I'm trying to solve". Ironically though, I think that because I've repeatedly tried and tested AI and it falls flat on its face over and over. However, this article makes me more hopeful that AI actually could be getting smarter.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#213

Earlier quoted context omitted.

Math and coding competition problems are easier to train because of strict rules and cheap verification. But once you go beyond that to less defined things such as code quality, where even humans have hard time putting down concrete axioms, they start to hallucinate more and become less useful. We are missing the value function that allowed AlphaGo to go from mid range player trained on human moves to superhuman by p…

Do we need all that if we can apply AI to solve practical problems today?

Depends on the cost.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#214
post #187

Earlier quoted context omitted.

> The verifier taught AlphaGo that move Ok so it sounds like you want to give the rules of Go credit for that move, lol.

It feels like you're purposefully ignoring the logical points OP gives and you just really really want to anthropomorphize AlphaGo and make us appreciate how smart it (should I say he/she?) is ... while no one is even criticising the model's capabilities, but analyzing it.

[dead]

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#215
post #176

Earlier quoted context omitted.

Are the only two options AI doubter and AI believer?

All I hear about are AI believers and AI-doubters-just-turned-believers

Hey, I'm a real person. Here's my website. I have YouTube videos up with my real name and face. https://validark.dev

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#216
post #55

I have long said I am an AI doubter until AI could print out the answers to hard problems or ones requiring tons of innovation. Assuming this is verified to be correct (not by AI) then I just became a believer. I would like to see a few more AI inventions to know for sure, but wow, it really is a new and exciting world. I really hope we use this intelligence resource to make the world better.

AI is a remixer; it remixes all known ideas together. It won't come up with new ideas though; the LLMs just predict the most likely next token based on the context. That means the group of characters it outputs must have been quite common in the past. It won't add a new group of characters it has never seen before on its own.

The ability for some people to perpetually move the goalpost will never cease to amaze me.

I guess that's one way to tell us apart from AIs.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#217
post #159

Earlier quoted context omitted.

It generated a proof that was close enough to something in its training data to be generated.

Do you know that from reading the proof, or are you just assuming this based on what you think LLMs should be capable of? If the latter, what evidence would be required for you to change your mind? - Edit: I can't reply, probably because the comment thread isn't allowed to go too deep, but this is a good argument. In my mind the argument isn't that coding is harder than math, but that the problems had resisted soluti…

1) this is a proof by example 2) the proof is conducted by writing a python program constructing hypergraphs 3) the consensus was this was low-hanging fruit ready to be picked, and tactics for this problem were available to the LLM

So really this is no different from generating any python program. There are also many examples of combinatoric construction in python training sets.

It's still a nice result, but it's not quite the breakthrough it's made out to be. I think that people somehow see math as a "harder" domain, and are therefore attributing more value to this. But this is a quite simple program in the end.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#218
post #143
post #63

Earlier quoted context omitted.

But human researchers are also remixers. Copying something I commented below: > Speaking as a researcher, the line between new ideas and existing knowledge is very blurry and maybe doesn't even exist. The vast majority of research papers get new results by combining existing ideas in novel ways. This process can lead to genuinely new ideas, because the results of a good project teach you unexpected things.

>But human researchers are also remixers. Some human researchers are also remixers to Some degree. Can you imagine AI coming up with refraction & separation lie Newton did?

That sets a vastly higher bar than what we're talking about here. You're comparing modern AI to one of the greatest geniuses in human history. Obviously AI is not there yet.

That being said, I think this is a great question. Did Einstein and Newton use a qualitatively different process of thought when they made their discoveries? Or were they just exceedingly good at what most scientists do? I honestly don't know. But if LLMs reach super-human abilities in math and science but don't make qualitative leaps of insight, then that could suggest that the answer is 'yes.'

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#219

Earlier quoted context omitted.

> I don't see this getting better. We went from 2 + 7 = 11 to "solved a frontier math problem" in 3 years, yet people don't think this will improve?

But can it count the R's in strawberry?

That question is equivalent to asking a human to add the wavelengths of those two colors and divide it by 3.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#220

Earlier quoted context omitted.

Software developers have spent decades at this point discounting and ignoring almost all objective metrics for software quality and the industry as a whole has developed a general disregard for any metric that isn't time-to-ship (and even there they will ignore faster alternatives in favor of hyped choices). (Edit: Yes, I'm aware a lot of people care about FP, "Clean Code", etc., but these are all red herrings that d…

But we don't even seem to be getting faster time-to-ship in any way that anybody can actually measure; it's always some vague sense of "we're so much more productive".

That's a fair observation and one that I don't really have an answer for. I can say from personal experience that I believe that shipping nonsense code has never been faster. That's just an anecdote, obviously.

We need a bigger version of the METR study on perceived vs. real productivity[0], I guess. It's a thankless job, though, since people will assume/state even at publication time that "Everything has progressed so much, those models and agents sucked, everything is 10 times better now!" and you basically have to start a new study, repeat ad infinitum.

One problem that really complicates things is that the net competency of these models seems really spotty and uneven. They're apparently out here solving math problems that seemingly "require thinking", but at the same time will write OpenGL code that will produce black screens on basically every driver, not produce the intended results and result in hours of debugging time for someone not familiar enough. That's despite OpenGL code being far more prevalent out there than math proofs, presumably. How do you reliably even theorize about things like this when something can be so bad and (apparently) so good at the same time?

0 - https://metr.org/blog/2025-07-10-early-2025-ai-experienced-o...

Post reply on HN