Live data from Hacker News

Epoch confirms GPT5.4 Pro solved a frontier math open problem

epoch.ai

531–540 of 744 posts

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#531

I don't know why I am still perpetually shocked that the default assumption is that humans are somehow unique. It's this pervasive belief that underlies so much discussion around what it means to be intelligent. The null hypothesis goes out the window. People constantly make comments like "well it's just trying a bunch of stuff until something works" and it seems that they do not pause for a moment to consider whethe…

It's only because humans came up with a problem, worked with the ai and verified the result that this achievement means anything at all. An ai "checking its own work" is practically irrelevant when they all seem to go back and forth on whether you need the car at the carwash to wash the car. Undoubtedly people have been passing this set of problems to ai's for months or years and have gotten back either incorrect res…

The only things moving faster than AI are the goalposts in conversations like this. Now we're at "sure, AI can solve novel problems, but it can't come up with the problems themselves on its own!"

I'm curious to see what the next goalpost position is.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#532

Earlier quoted context omitted.

This is too dependent on what you mean by "unique", though. What do we have that apes don't, and which directly enables intelligence? What do we have that LLMs don't? What do LLMs have that we don't? I don't think we know enough to definitively say "it's this bit that gives us intelligence, and there's no way to have intelligence without it". We just see what we have, and what animals lack, and we say "well it's prob…

> What do we have that apes don't, and which directly enables intelligence? Again, there are multiple fields of study with tons of amazingly detailed answers to this. We know about specific proteins, specific brain structures, we know about specific cognitive capabilities in the abstract, etc. > What do we have that LLMs don't? Again, quite a lot is already known about this. This feels a bit like you're starting to e…

It'd probably be more productive for you to actually back up your claims with these things we know from neuroscience, rather than just stating that we know things, and so therefore you're right. What do we know?

EDIT: can't reply, so I'll just update here:

You're arguing that the mechanism that produces human intelligence is unique, so therefore the intelligence itself is somehow fundamentally different from the intelligence an LLM can produce. You haven't shown that, you just keep saying we know it's true. How do we know?

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#533

Earlier quoted context omitted.

Logical fallacies are vastly overrated. Unless the conversation is formal logic in the first place, "logical fallacies" are just a way to apply quick pattern matching to dismiss people without spending time on more substantive responses. In this case, both you and the other are speculating about the near future of a thing, neither of you knows.

Hard to make a more substantive response when the OP’s entire comment was a one-sentence logical fallacy. I’m not cherry-picking here. > In this case, both you and the other are speculating about the near future of a thing, neither of you knows. One of us is making a much grander claim than the other: - LLMs have limitless potential for growth; because they are not capable of something today does not mean they won’t…

The post you replied to was:

> We went from 2 + 7 = 11 to "solved a frontier math problem" in 3 years, yet people don't think this will improve?

All that says is that the speaker thinks models will improve past where they are today. Not that it's a logical certainty (the first thing you jumped on them for), and certainly not anything about "limitless potential for growth" (which nobody even mentioned). With replies like this, invoking fallacies and attacking claims nobody made, you're adding a lot of heat and very little light here (and a few other threads on the page).

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#534
post #479

Earlier quoted context omitted.

How confident are you that this knowledge was not part of the training data? Was there no stackoverflow questions/replies with it, no tech forum posts, private knowledge bases, etc? Not trying to diminish its results, just one should always assume that LLMs have a rough memory on pretty much the whole of the internet/human knowledge. Google itself was very impressive back then in how it managed to dig out stuff inter…

Do you think AlphaGo is regurgitating human gameplay? No it’s not: it’s learning an optimal policy based on self play. That is essentially what you’re seeing with agents. People have a very misguided understanding of the training process and the implications of RL in verifiable domains. That’s why coding agents will certainly reach superhuman performance. Straw/steel man depending on what you believe: “But they won’t…

How does alphago come into picture? It works in a completely different way all together.

I'm not saying that LLMs can't solve new-ish problems, not part of the training data, but they sure as hell not got some Apple-specific library call from a divine revelation.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#535

Earlier quoted context omitted.

I think "novel" is ill defined here, perhaps. LLMs do appear to be poor general reasoners[0], and it's unclear if they'll improve here. It would be unintuitive for them to be good at this, given that we know exactly how they're implemented - by looking at text and then building a statistical model to predict the next token. From this, if we wanted to commit to LLMs having generalizable knowledge, we'd have to assume…

The “good deal of evidence” is everywhere. The proof is in the pudding. Of course you can find failure modes, the blog article (not an actual paper?) rightfully derides benchmarks and then…creates a benchmark? Designed to elicit failure modes, ok so what? As if this is surprising to anyone and somehow negates everything else? Anyone who says that “statistical models for next token generation” are unlikely to provide…

> The “good deal of evidence” is everywhere. The proof is in the pudding.

I'm open! Please, by all means.

> the blog article (not an actual paper?) rightfully derides benchmarks and then…creates a benchmark?

The blog article is a review of benchmarking methodologies and the issues involved by a PhD neuroscientist who works directly on large language models and their applications to neuroscience and cognition, it's probably worth some consideration.

> Anyone who says that “statistical models for next token generation” are unlikely to provide emergent intelligence I think is really not understanding what a statistical model for next token generation really means.

Okay.

> That is a proxy task DESIGNED to elicit intelligence because in order to excel at that task beyond a certain point you need to develop the right abstractions and decide how to manipulate them to predict the next token (which, by the way, is only one of many many stages of training).

This isn't a great argument. It seems to say that in order for LLMs to do well they must have emergent intelligence. That is not evidence for LLMs having emergent intelligence, it's just stating that a goal would be to have it.

As I said, a theoretical framework with real tests would be great. That's how science is done, I don't really think I'm asking for a lot here?

> It’s like saying “I think it’s surprising that a jumble of trillions of little cells zapping each other would produce emergent intelligence” while ignoring the fact that brains are clearly intelligent.

Well, it is a bit surprising. But we have an extremely robust model for exactly that - there are fields dedicated to it, we can create simulations and models, we can perform interventative analysis, we have a theory and falsifying test cases, etc. We don't just say "clearly brains are intelligent, therefor intelligence is an emergent property of cells zapping" lol that would be absurd.

So I'm just asking for you to provide a model and evidence. How else should I form my beliefs? As I've expressed, I have reasons to find the idea of emergent logic from statistical models surprising, and I have no compelling theory to account for that nor evidence to support that. If you have a theory and evidence, provide it! I'd be super interested, I'm in no way ideologically opposed to the idea. I'm a functionalist so I fundamentally believe that we can build intelligent systems, I'm just not convinced that LLMs are doing that - I'm not far though, so please, what's the theory?

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#536
post #473

Earlier quoted context omitted.

It probably can, but won't realize that and it won't be efficient in that. LLM can shuffle tokens for an enormous number of tries and eventually come up with something super impressive, though as you yourself have mentioned, we would need to have a mandatory verification loop, to filter slop from good output and how to do it outside of some limited areas is a big question. But assuming we have these verification loop…

How can people look at - clear generalizability - insane growth rates (go back and look at where we were maybe 2 years ago and then consider the already signed compute infrastructure deals coming online) And still say with a straight face that this is some kind of parlor trick or monkeys with typewriters. we don’t need to run LLMs for years. The point is look at where we are today and consider performance gets 10x ch…

I was talking about highest difficulty problems only, in the scope of that comment. Sure at mundane tasks they are useful and we optimizing that constantly.

But for super hard tasks, there is no situation when you just dump a few papers for context add a prompt and LLM will spit out correct answer. It's likely that a lead on such project would need to additionally train LLM on their local dataset, then parse through a lot of experimental data, then likely run multiple LLMs for for many iterations homing on the solution, verifying intermediate results, then repeating cycle again and again. And in parallel the same would do other team members. All in all, for such a huge hard task a year of cumulative machine-hours is not something outlandish.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#537
post #534

Earlier quoted context omitted.

Do you think AlphaGo is regurgitating human gameplay? No it’s not: it’s learning an optimal policy based on self play. That is essentially what you’re seeing with agents. People have a very misguided understanding of the training process and the implications of RL in verifiable domains. That’s why coding agents will certainly reach superhuman performance. Straw/steel man depending on what you believe: “But they won’t…

How does alphago come into picture? It works in a completely different way all together. I'm not saying that LLMs can't solve new-ish problems, not part of the training data, but they sure as hell not got some Apple-specific library call from a divine revelation.

AlphaGo comes into the picture to explain that in fact coding agents in verifiable domains are absolutely trained in very similar ways.

It’s not magic they can’t access information that’s not available but they are not regurgitating or interpolating training data. That’s not what I’m saying. I’m saying: there is a misconception stemming from a limited understanding of how coding agents are trained that they somehow are limited by what’s in the training data or poorly interpolating that space. This may be true for some domains but not for coding or mathematics. AlphaGo is the right mental model here: RL in verifiable domains means your gradient steps are taking you in directions that are not limited by the quality or content of the training data that is used only because starting from scratch using RL is very inefficient. Human training data gives the models a more efficient starting point for RL.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#538

Earlier quoted context omitted.

> What do we have that apes don't, and which directly enables intelligence? Again, there are multiple fields of study with tons of amazingly detailed answers to this. We know about specific proteins, specific brain structures, we know about specific cognitive capabilities in the abstract, etc. > What do we have that LLMs don't? Again, quite a lot is already known about this. This feels a bit like you're starting to e…

It'd probably be more productive for you to actually back up your claims with these things we know from neuroscience, rather than just stating that we know things, and so therefore you're right. What do we know? EDIT: can't reply, so I'll just update here: You're arguing that the mechanism that produces human intelligence is unique, so therefore the intelligence itself is somehow fundamentally different from the inte…

I don't need to do that unless you think that neurons interact exactly the way that LLMs do? That said, we have detailed, microscopic models of neurons, the ability to even simulate brain activity, intervention studies where we can make predictions, interact with brains in various ways, and then validate against predictions, we have cognitive benchmarks that we can apply to different animals or animals in different stages of development that we can then tie to specific brain states and brain development, etc.

So we're in a very good position to say quite a lot about the brain, an incredible amount really. And that puts us in a very good position to say that our brain is very different from other animal brains, and certainly in a very good position to say that's very different from an LLM.

Now, you can argue that an LLM is functionally equivalent to the brain, but given that it's so structurally distinct, and seemingly functions in a radically different way due to the nature of that structure, I'd put it on you to draw symmetries and provide evidence of that symmetry.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#539

I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…

Ok, I'll bite. Show me an LLM that comes up with a new math operator. Or which will come up with theory of relativity if only Newton physics is in its training dataset. That it could remix existing ideas which leads to novel insights is expected, however the current LLMs can't come up with paradigm shifts that require novel insights. Even humans have a rather limited time they can come up with novel insights (when they are young, capable of latent thinking, not yet ossified from the existing formalization of science and their brain is still energetically capable without vascular and mitochondrial dysfunction common as we age).

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#540
post #229

Earlier quoted context omitted.

Maybe to get a real breakthrough we have to make programming languages / tools better suited for LLM strengths not fuss so much about making it write code we like. What we need is correct code not nice looking code.

> programming languages / tools better suited for LLM strengths The bitter lesson is that the best languages / tools are the ones for which the most quality training data exists, and that's pretty much necessarily the same languages / tools most commonly used by humans. > Correct code not nice looking code "Nice looking" is subjective, but simple, clear, readable code is just as important as ever for projects to be l…

>> simple, clear, readable code is just as important as ever for projects to be long-term successful

Is it though? I'm a long-time code purist, but I am beginning to wonder about the assumptions underlying our vocation.

Post reply on HN