Live data from Hacker News

Epoch confirms GPT5.4 Pro solved a frontier math open problem

epoch.ai

551–560 of 744 posts

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#551

Earlier quoted context omitted.

It's only because humans came up with a problem, worked with the ai and verified the result that this achievement means anything at all. An ai "checking its own work" is practically irrelevant when they all seem to go back and forth on whether you need the car at the carwash to wash the car. Undoubtedly people have been passing this set of problems to ai's for months or years and have gotten back either incorrect res…

The only things moving faster than AI are the goalposts in conversations like this. Now we're at "sure, AI can solve novel problems, but it can't come up with the problems themselves on its own!" I'm curious to see what the next goalpost position is.

> I'm curious to see what the next goalpost position is.

I am as well. That's the point. Ai can do some things well and other things better than humans, but so can a garden hose and all technology. Is ai just a tool or is it the future of all work? By setting goalposts we can see whether or not it is living up to the hype that we're collectively spending trillions on.

The garden hose manufacturers aren't claiming that they're going to replace all human workers, so we don't set those kinds of goalposts to measure whether it's doing that.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#552
post #55

I have long said I am an AI doubter until AI could print out the answers to hard problems or ones requiring tons of innovation. Assuming this is verified to be correct (not by AI) then I just became a believer. I would like to see a few more AI inventions to know for sure, but wow, it really is a new and exciting world. I really hope we use this intelligence resource to make the world better.

AI is a remixer; it remixes all known ideas together. It won't come up with new ideas though; the LLMs just predict the most likely next token based on the context. That means the group of characters it outputs must have been quite common in the past. It won't add a new group of characters it has never seen before on its own.

Move 37.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#553

Earlier quoted context omitted.

It'd probably be more productive for you to actually back up your claims with these things we know from neuroscience, rather than just stating that we know things, and so therefore you're right. What do we know? EDIT: can't reply, so I'll just update here: You're arguing that the mechanism that produces human intelligence is unique, so therefore the intelligence itself is somehow fundamentally different from the inte…

I don't need to do that unless you think that neurons interact exactly the way that LLMs do? That said, we have detailed, microscopic models of neurons, the ability to even simulate brain activity, intervention studies where we can make predictions, interact with brains in various ways, and then validate against predictions, we have cognitive benchmarks that we can apply to different animals or animals in different s…

I'm following this mini-thread with interest but I've arrived here and I confess, I don't really know what your argument is.

I think this all stems from you objecting to this statement:

"I don't know why I am still perpetually shocked that the default assumption is that humans are somehow unique."

I think you're being uncharitable in how you interpret that. Human's are unique in the most literal reading of this sentence, we don't have anything else like humans. But the context is the ability to reason and people denying that a machine is reasoning, even though it looks like reasoning.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#554
post #454

Earlier quoted context omitted.

"Truly novel" is fast becoming a True Scotsman.

No True Novelty, No True Understanding, etc. The problem with these bromides is not that they're wrong, it's that they're not even wrong. They're predictive nulls. What observable differences can we expect between an entity with True Understanding and an entity without True Understanding? It's a theological question, not a scientific one. I'm not an AI booster by any means, but I do strongly prefer we address the que…

Agreed. We should be asking what the machines measurably can or can't do. If it can't be measured, then it doesn't matter from an engineering standpoint. Does it have a soul? Can't measure it, so it doesn't matter.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#555

Earlier quoted context omitted.

I don't need to do that unless you think that neurons interact exactly the way that LLMs do? That said, we have detailed, microscopic models of neurons, the ability to even simulate brain activity, intervention studies where we can make predictions, interact with brains in various ways, and then validate against predictions, we have cognitive benchmarks that we can apply to different animals or animals in different s…

I'm following this mini-thread with interest but I've arrived here and I confess, I don't really know what your argument is. I think this all stems from you objecting to this statement: "I don't know why I am still perpetually shocked that the default assumption is that humans are somehow unique." I think you're being uncharitable in how you interpret that. Human's are unique in the most literal reading of this sente…

They're shocked that people believe that humans are unique. I explained why that shouldn't be shocking. I think I was pretty charitable here, I gave an alternative option for what they could mean in my very first reply:

> Unless you mean "fundamentally unique" in some way that would persist - like "nothing could ever do what humans do".

> I don't really know what your argument is.

I just said that I think that we have very good reasons for believing that human cognition is unique. The response was seemingly that we don't have enough of an understanding of intelligence to make that judgment. I've stated that I think we do have enough of an understanding of intelligence to make that judgment, and I've appealed to the many advances in relevant feilds.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#556
post #507

Earlier quoted context omitted.

We have a tremendous amount of raw information flowing through our brains 24/7 from before we are born, from the external world through all our senses and from within our minds as it attempts to make sense of that information, make predictions, generally reason about our existence, hallucinate alternative realities, etc. etc. If you were able to somehow capture all that information in full detail as you've had access…

Humans are "multi-modal". Sure we get plenty of non-textual information, but LLMs were trained on basically every human-written world ever. They definitely see many orders of magnitude more language than any human has ever seen. And yet humans get fluent based after 3+ years.

For sure, it seems like there's something there primed to pick up human language quickly, clearly evolutionarily driven.

Not necessarily so for the dynamics of magnetic fields, or nonhuman animal communications, or dark energy/matter.

We are bombarded nonstop by magnetic fields, nonhuman animal communications, and live in a universe which seems to be majority dominated by dark energy and matter, and yet understand little to none of it all.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#557

Earlier quoted context omitted.

The “good deal of evidence” is everywhere. The proof is in the pudding. Of course you can find failure modes, the blog article (not an actual paper?) rightfully derides benchmarks and then…creates a benchmark? Designed to elicit failure modes, ok so what? As if this is surprising to anyone and somehow negates everything else? Anyone who says that “statistical models for next token generation” are unlikely to provide…

> The “good deal of evidence” is everywhere. The proof is in the pudding. I'm open! Please, by all means. > the blog article (not an actual paper?) rightfully derides benchmarks and then…creates a benchmark? The blog article is a review of benchmarking methodologies and the issues involved by a PhD neuroscientist who works directly on large language models and their applications to neuroscience and cognition, it's pr…

> The “good deal of evidence” is everywhere. The proof is in the pudding. I'm open! Please, by all means.

sure here are but a few: [1] you get smooth gains in reasoning with more RL train-time compute and more test-time compute (o1)

[2] DeepSeek-R1 showed that RL on verifiable rewards produces behavior like backtracking, adaptation, reflection, etc.

[3] SWE-Bench is a relatively decent benchmark and perf here is continually improving — these are real GitHub issues in real repos

[4] MathArena — still good perf on uncontaminated 2025 AIME problems

[5] the entire field of reinforcement learning, plus successes in other fields with verifiable domains (e.g. AlphaGo); Bellman updates will give you optimal policies eventually

[6] Anthropics cool work looking effectively at biology of a large language models: https://transformer-circuits.pub/2025/attribution-graphs/met... — if you trace internal circuits in Haiku 3.5 you see what you expect from a real reasoning system: planning ahead, using intermediate concepts, operating in a conceptual latent space (above tokens). And thats Haiku 3.5!!! We’re on Opus 4.6 now…

people like to move goalposts whenever a new result comes out, which is silly. Could AI systems do this 2 years ago? No. I don’t know how people don’t look at robust trends in performance improvement, combined with verifiable RL rewards, and can’t understand where things are going.

> The blog article is a review of benchmarking methodologies and the issues involved by a PhD neuroscientist who works directly on large language models and their applications to neuroscience and cognition, it's probably worth some consideration.

Appeals to authority are a fine prior, but lo and behold I also have a PhD and have worked on and led benchmark development professionally for several years at an AI lab. That’s ultimately no reason to really trust either of us. As I said, the blog post rightfully decries benchmarks but it then presents a new benchmark as though that isn’t subject to all of the same problems. It’s a good article! I think they do a good job here! I agree with all of their complaints about benchmarks! It rightfully identifies failure modes, and there are plenty of other papers pointing out similar failure modes. Reasoning is still brittle, lots of areas where LLMs/agentic systems fail in ways that are incredible given their talent in other areas. But you pretend as though this is definitive evidence that “LLMs are poor general reasoners”. This is just not true, but it is true that they are brittle and fallible in weird ways, today.

> This isn't a great argument. It seems to say that in order for LLMs to do well they must have emergent intelligence. That is not evidence for LLMs having emergent intelligence, it's just stating that a goal would be to have it.

"They do well, therefore intelligence" is not an argument, sure. But that’s also not what I’m saying. The Occam’s razor here is that reasoning-like computation is the best explanation for an increasing amount of the observed behavior, especially in fresh math and real software tasks where memorization is a much worse fit.

> As I said, a theoretical framework with real tests would be great. That's how science is done, I don't really think I'm asking for a lot here?

I would encourage you to read Kuhn’s structure of scientific revolutions. "That’s how science is done" is a bit of an oversimplification of how the sausage is made here. Real science moves forward in a messy mix of partial theory + better measurements + interventions long before anyone has some sort of grand unified framework. Neuroscience is no different here. And I would say at this point with LLMs we now do have pretty decent tests: fresh verifiable-task evals, mechanistic circuit tracing, causal activation patching, and scaling results for RL/test-time compute. The claim that there is no framework + no real tests is just not true anymore. It’s not like we have some finished theory of reasoning, but thats a bit of an unfair demand at this point and is asymmetrical as well.

> It’s like saying “I think it’s surprising that a jumble of trillions of little cells zapping each other would produce emergent intelligence” while ignoring the fact that brains are clearly intelligent.

>> Well, it is a bit surprising. But we have an extremely robust model for exactly that - there are fields dedicated to it, we can create simulations and models, we can perform interventative analysis, we have a theory and falsifying test cases, etc. We don't just say "clearly brains are intelligent, therefor intelligence is an emergent property of cells zapping" lol that would be absurd.

>> So I'm just asking for you to provide a model and evidence. How else should I form my beliefs? As I've expressed, I have reasons to find the idea of emergent logic from statistical models surprising, and I have no compelling theory to account for that nor evidence to support that. If you have a theory and evidence, provide it! I'd be super interested, I'm in no way ideologically opposed to the idea. I'm a functionalist so I fundamentally believe that we can build intelligent systems, I'm just not convinced that LLMs are doing that - I'm not far though, so please, what's the theory?

The model is: reasoning is not inherently human, it’s mathematical. It falls easily within the purview of RL, statistics, representation, optimization, etc, and to claim otherwise would require evidence.

What is the robust model for reasoning in humans again? Simulations and models — what are these? Interventative analysis — we can’t do this with LLMs? Falsifying test cases — what would satisfy you here beyond everything I’ve presented above? Also I’m confused by your last part. You say “brains are intelligent” ==> “intelligence is an emergent property of cells zapping” is absurd, but why? You start from the position that brains are intelligent, so why is this absurd within your argument? Brains _are_ made up of real, physical atoms organized into molecules organized into cells organized into a coordinated system, and…that’s it? What’s missing here?

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#558

Earlier quoted context omitted.

I’ve seen this style of take so much that I’m dying for someone to name a logical fallacy for it, like “appeal to progress” or something. Step away from LLMs for a second and recognize that “Yesterday it was X, so today it must be X+1” is such a naive take and obviously something that humans so easily fall into a trap of believing (see: flying cars).

In finance we say "past performance does not guarantee future returns." Not because we don't believe that, statistically, returns will continue to grow at x rate, but because there is a chance that they won't. The reality bias is actually in favour of these getting better faster, but there is a chance they do not.

this is true because markets are generally efficient. It's very hard to find predictive signals. That is a completely different space than what we're talking about here. Performance is incredibly predictable through scaling laws that continue to hold even at the largest scales we've built

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#559

Earlier quoted context omitted.

> I don't see this getting better. We went from 2 + 7 = 11 to "solved a frontier math problem" in 3 years, yet people don't think this will improve?

I’ve seen this style of take so much that I’m dying for someone to name a logical fallacy for it, like “appeal to progress” or something. Step away from LLMs for a second and recognize that “Yesterday it was X, so today it must be X+1” is such a naive take and obviously something that humans so easily fall into a trap of believing (see: flying cars).

Hmm...the sun comes up today is a pretty good bet that the sun comes up tomorrow.

We have robust scaling laws that continue to hold at the largest scales. It is absolutely a very safe bet that more compute + more training + algorithmic improvements will certainly improve performance it's not like we're rolling a 1 trillion dollar die.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#560
post #97

I have long said I am an AI doubter until AI could print out the answers to hard problems or ones requiring tons of innovation. Assuming this is verified to be correct (not by AI) then I just became a believer. I would like to see a few more AI inventions to know for sure, but wow, it really is a new and exciting world. I really hope we use this intelligence resource to make the world better.

It 100% will not be used to make the world better and we all know it will be weaponised first to kill humans like all preceding tech

Most tech gets used for good and bad.
Post reply on HN