Live data from Hacker News

Epoch confirms GPT5.4 Pro solved a frontier math open problem

epoch.ai

681–690 of 744 posts

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#681

Earlier quoted context omitted.

This is special pleading. Long multiplication is a trivial form of reasoning that is taught at elementary level. Furthermore, the LLM isn't doing things "in its head" - the headline feature of GPT LLMs is attention across all previous tokens, all of its "thoughts" are on paper. That was Opus with extended reasoning, it had all the opportunity to get it right, but didn't. There are people who can quickly multiply such…

Furthermore, the LLM isn't doing things "in its head" - the headline feature of GPT LLMs is attention across all previous tokens, all of its "thoughts" are on paper LOL, talk about special pleading. Whatever it takes to reshape the argument into one you can win, I guess... LLMs don't reason. Let's see you do that multiplication in your head. Then, when you fail, we'll conclude you don't reason. Sound fair?

I can do it with a scratch pad. And I can also tell you when the calculation exceeds what I can do in my head and when I need a scratch pad. I can also check a long multiplication answer in my head (casting 9s, last digit etc.) and tell if there’s a mistake.

The LLMs also have access to a scratch pad. And importantly don’t know when they need to use it (as in, they will sometimes get long multiplication right if you ask them to show their work but if you don’t ask them to they will almost certainly get it wrong).

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#682

Earlier quoted context omitted.

This is special pleading. Long multiplication is a trivial form of reasoning that is taught at elementary level. Furthermore, the LLM isn't doing things "in its head" - the headline feature of GPT LLMs is attention across all previous tokens, all of its "thoughts" are on paper. That was Opus with extended reasoning, it had all the opportunity to get it right, but didn't. There are people who can quickly multiply such…

Furthermore, the LLM isn't doing things "in its head" - the headline feature of GPT LLMs is attention across all previous tokens, all of its "thoughts" are on paper LOL, talk about special pleading. Whatever it takes to reshape the argument into one you can win, I guess... LLMs don't reason. Let's see you do that multiplication in your head. Then, when you fail, we'll conclude you don't reason. Sound fair?

The conclusion that LLMs don't reason is not a consequence of them not being able to do arithmetic, so your argument isn't valid.

Also, see https://news.ycombinator.com/newsguidelines.html

"Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.

Comments should get more thoughtful and substantive, not less, as a topic gets more divisive.

When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."

Don't be curmudgeonly. Thoughtful criticism is fine, but please don't be rigidly or generically negative."

etc.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#683
post #677
post #673

Earlier quoted context omitted.

Searle's Chinese Room is a fallacious mess ... see the works of Larry Hauser, e.g., https://philpapers.org/rec/HAUNGT and https://philpapers.org/rec/HAUSCB-2 The importance of Searle's Chinese Room is how such extraordinarily bad argumentation has persuaded so many people open to it. And the literature about philosophical zombies is contentious, to say the least, and much of it is also among the worst arguments in ph…

I would hope that philosophy would be exempt from accusations of arguments from authority. I say I don’t want to fight exactly because I don’t want to come off like a jerk because I’m arguing. If the Chinese Room is a mess, I welcome the argument, and will happily read the paper. I’m less open to push back against philosophical zombies, as the argument seems trivially plausible, from a position of solipsism.

Philosophy may be exempt from accusations of arguments from authority--because that's a category mistake--but philosophers certainly aren't.

Hauser's papers are just a part of a large literature rejecting/refuting Searle's Chinese Room, but he has probably taken Searle more seriously than most. After Searle's well known response that waves away numerous objections, many people dismissed him as acting in bad faith. (It would have been even worse if they had known about the accusations of sexual assault. Sure, that would be ad hominem and intellectually dishonest, but we're talking about human beings, same as with arguments from authority.) See, e.g., https://www.nybooks.com/articles/1995/12/21/the-mystery-of-c... where Daniel Dennett writes:

> For his part, he has one argument, the Chinese Room, and he has been trotting it out, basically unchanged, for fifteen years. It has proven to be an amazingly popular number among the non-experts, in spite of the fact that just about everyone who knows anything about the field dismissed it long ago. It is full of well-concealed fallacies. By Searle’s own count, there are over a hundred published attacks on it. He can count them, but I guess he can’t read them, for in all those years he has never to my knowledge responded in detail to the dozens of devastating criticisms they contain; he has just presented the basic thought experiment over and over again. I just went back and counted: I am dismayed to discover that no less than seven of those published criticisms are by me (in 1980, 1982, 1984, 1985, 1987, 1990, 1991, 1993).

etc. If you've never read any of this literature yet can facilely write what you did above about Searle's discussion of the Chinese Room being "the most important work here", I don't expect you to start now ... but at least reconsider posing as a philosopher who is knowledgeable about such things.

Your reason to be less open to "push back against" (an odd formulation--the burden is on those who claim that they are conceivable, and therefore physicalism is false) philosophical zombies seems to hinge on another radical failure to understand the issue and unfamiliarity with the literature.

Philosophical zombies are completely independent of solipsism. The conceivability of zombies says that, if this is a world in which you are the sole inhabitant and you are conscious, then there is a possible world that is physically identical to this world and has the same physical laws, but the sole inhabitant (scoofy'), while physically identical to you and behaves identically, isn't conscious. That is, consciousness is not a consequence of physical laws and contingencies but is some sort of ethereal goop that accompanies physical entities. Of course Chalmers and other modern dualists don't subscribe to Descartes' substance dualism, but their attempts to formulate "process dualism" or some other nonsense solely because they need some alternative to physicalism--which they reject because they are hopelessly confused about the nature of consciousness and "qualia"--are frankly incoherent.

Maybe read Kirk's book and learn something about the subject. Here's a review that gives you a peek at what you'll find there: https://view.officeapps.live.com/op/view.aspx?src=https%3A%2...

Over and out.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#684
post #665

Earlier quoted context omitted.

This is special pleading. Long multiplication is a trivial form of reasoning that is taught at elementary level. Furthermore, the LLM isn't doing things "in its head" - the headline feature of GPT LLMs is attention across all previous tokens, all of its "thoughts" are on paper. That was Opus with extended reasoning, it had all the opportunity to get it right, but didn't. There are people who can quickly multiply such…

Mathematics is not the only kind of reasoning, so your conclusion is false. The human brain also has compartments for different types of activities. Why shouldn't an AI be able to use tools to augment its intelligence?

I used the mathematics example only because the GP did. There are many other examples of non-reasoning, including some papers (as recent as Feb).

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#685

Earlier quoted context omitted.

LLMs are notoriously terrible at multiplying large numbers: https://claude.ai/share/538f7dca-1c4e-4b51-b887-8eaaf7e6c7d3 > Let me calculate that. 729,278,429 × 2,969,842,939 = 2,165,878,555,365,498,631 Real answer is: https://www.wolframalpha.com/input?i=729278429*2969842939 > 2 165 842 392 930 662 831 Your example seems short enough to not pose a problem.

I thought it might do better if I asked it to do long-form multiplication specifically rather than trying to vomit out an answer without any intermediate tokens. But surprisingly, I found it doesn't do much better.

Other comments indicate that asking it to do long multiplication does work, but the varying results makes sense: LLMs are probabilistic, you probably rolled an unlikely result.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#686

Earlier quoted context omitted.

LLMs are notoriously terrible at multiplying large numbers: https://claude.ai/share/538f7dca-1c4e-4b51-b887-8eaaf7e6c7d3 > Let me calculate that. 729,278,429 × 2,969,842,939 = 2,165,878,555,365,498,631 Real answer is: https://www.wolframalpha.com/input?i=729278429*2969842939 > 2 165 842 392 930 662 831 Your example seems short enough to not pose a problem.

This doesn’t address the author’s point about novelty at all. You don’t need 100% accuracy to have the capability to solve novel problems.

It does address the GP comment about math.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#687
post #581

Earlier quoted context omitted.

If all art is derivative then the earlier statement is a tautology. People still call things other people do novel. There's clear social proof that humans do things that other humans consider novel. Otherwise the word would probably not exist. Just today I wrote a python program that did not resemble anything I'd written before, nor had I seen anything similar. I had to reason it out myself. That passes thr test that…

Your threshold for "resemble" is obviously quite high, which is fair, but assuming that you're an encultured programmer your python code represents other people's python code. It might be doing something novel, but that thing it's doing is interacting or in response to, or otherwise relative to existing concepts you learned or saw elsewhere. All art is derivative, we can do things other people haven't done before but…

I'm not talking about wacky. My barrier for novel is 1) new capabilities 2) useful, and 3) end-to-end tested.

For example, what I refered to that I've written is a dynamic storage solution for n-dimensional grids, that can grow arbitrarily in any direction, and is locally dense (organized into spatially indexed blocks of contiguous data).

I had never considered this problem before, and I certainly had never seen a solution before (even though there may well be one).

I worked it out on paper, considering how integer lattices can be partitioned and indexed, and then I transformed that into a design which I then implemented. Working purely from the design, not considering existing solutions.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#688
post #682

Earlier quoted context omitted.

Furthermore, the LLM isn't doing things "in its head" - the headline feature of GPT LLMs is attention across all previous tokens, all of its "thoughts" are on paper LOL, talk about special pleading. Whatever it takes to reshape the argument into one you can win, I guess... LLMs don't reason. Let's see you do that multiplication in your head. Then, when you fail, we'll conclude you don't reason. Sound fair?

The conclusion that LLMs don't reason is not a consequence of them not being able to do arithmetic, so your argument isn't valid. Also, see https://news.ycombinator.com/newsguidelines.html "Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes. Comments should get more thoughtful and substantive, not less, as a topic gets more divisive. When disagreeing, please reply to the argument inste…

Plenty of humans can't do arithmetic. Can they also not reason.

Reasoning isn't a binary switch. It's a multidimensional continuum. AI can clearly reason to some extent even if it also clearly doesn't reason in the same way that a human would.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#689

Earlier quoted context omitted.

In finance we say "past performance does not guarantee future returns." Not because we don't believe that, statistically, returns will continue to grow at x rate, but because there is a chance that they won't. The reality bias is actually in favour of these getting better faster, but there is a chance they do not.

this is true because markets are generally efficient. It's very hard to find predictive signals. That is a completely different space than what we're talking about here. Performance is incredibly predictable through scaling laws that continue to hold even at the largest scales we've built

I agree this is a new space and prediction volatility is much higher. We have evidence going back to at least 2019 that improvements have been exponential (https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com...). The benchmarks are all over the place because improvements don't happen in a straight line. Even composites aren't that useful because the last 10% improvement can require more effort than the first 90%.

To be frank, from what I can see, even if all progress stopped right now, it would take 1-2 decades to fully operationalise the existing potential of LLMs. There would be massive economic and social change. But progress is not stopping, and in some measurements, continues to improve exponentially. I really think this is incredibly transformative. Moreso than anything humanity has ever experienced. In the last year, OpenAI and potentially Claude have been working on recursive self-improvement. Meaning these models are designing better versions of themselves. This means we have effectively entered the singularity.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#690
post #600

Earlier quoted context omitted.

I studied philosophy focusing on the analytic school and proto-computer science. LLMs are going to force many people start getting a better understanding about what "Knowledge" and "Truth" are, especially the distinction between deductive and inductive knowledge. Math is a perfect field for machine learning to thrive because theoretically, all the information ever needed is tied up in the axioms. In the empirical wor…

> Math is a perfect field for machine learning to thrive because theoretically, all the information ever needed is tied up in the axioms. Not really; the normal way that math progresses, just like everything else, is that you get some interesting results, and then you develop the theoretical framework. We didn't receive the axioms; we developed them from the results that we use them to prove.

Axioms are, again, by definition, arbitrary. It is effectively irrelevant that we try to develop axioms so that the framework mirror the real world. Everything falls out of the axioms, period.

If you want to change the axioms to better reflect some aspect about life, that's all well and good, but everything will still fall out of the new axioms.

Post reply on HN