Live data from Hacker News

Epoch confirms GPT5.4 Pro solved a frontier math open problem

epoch.ai

591–600 of 744 posts

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#591

Earlier quoted context omitted.

Funding a few PhDs for a year costs orders of magnitude more than it did to solve this problem in inference costs. Also, this has been active research for some time. Or I guess the people working on it are just not as good as a random bunch of students? It's amazing the lengths that people go to maintain their worldview, even if it means belittling hardworking people. I take it you're not a mathematician. This is an…

Inference costs are heavily subsidised. My point was that we've spent trillions collectively on ai, and so far we have a few new proofs. It's been active research but the problem estimates only 5-10 people are even aware that it is a problem. I wrote "math phd's" not "random students", but regardless, I wouldn't know how you interpreted my statement that people could have discovered without ai this as "belittling the…

> You seem like a stupid person

And now you're belittling me. Yeah, good one, that'll convince people.

> out of control chatbot that can't comprehend basic arguments

I don't see how it is out of control. It is a tool. It is being used for a job. For low-level jobs it often succeeds. For tougher jobs, it is succeeding sufficiently often to be interesting. I don't care if it understands worldview semantics, that's for humans to do.

> we've spent trillions collectively on ai

The economics around AI do not suggest that continuing to perform large training runs is sustainable. That's also not relevant to the discussion. Once the training is done, further costs are purely on inference, and that is the comparison I was making.

> Inference costs are heavily subsidised

Even if you pay to run inference on your own hardware, economics of scale dictate that it is still cheaper than students.

> It's been active research but the problem estimates only 5-10 people are even aware that it is a problem.

That sounds about right for most pure math problems. Were you expecting more?

Let's not pretend that society would have invested that kind of money into pure mathematics research. It is extraordinarily difficult to get funding for that kind of work in most parts of the world. Mathematicians are relatively cheap, yes, but the money coming into AI was from blind VCs with a sense of grandeur. It wasn't to do maths research. If it's here anyway, and causing nightmares for actually teaching new students, may as well try to make some good of it. It has only recently crossed the edge of being useful. Most researchers I know are only now starting to consider it, mostly as a search engine, but some for proof assistance. Experiences a year ago were highly negative. They're a lot more positive now.

I'm trying to give a perspective from someone who actually does do math research at a senior level, who actually does have a half dozen math PhD students to supervise, to say that your blind attitude toward this is not sensible or helpful. Your comments about the problem being trivial do belittle the actual effort people have put into the problem without success. If they could easily have discovered this without AI, they would have already done so. Researchers do not have unlimited time and there are many more problems than students, especially good ones (hence my random comment).

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#592
post #578

Earlier quoted context omitted.

> Funding a few PhDs for a year costs orders of magnitude more than it did to solve this problem in inference costs. I don't think PhD students are sitting around and solving one problem for a year. Also PhD students are way cheaper

How many math PhD students do you have? If you set the problem right, something like this per year on average is a good pace. How are they cheaper? Your average grant where I am can pay for a couple of PhD students. I could afford to pay for inference costs out of my own salary, no grant needed. Completely different economic scales here. I like students better of course, but funding is drying up these days.

I was saying generally. I don't work in maths. PhD students do lots of other things than research. If we ask a PhD student to just solve these kinds of problems and nothing else, the student would do it without much difficulty.

I guess it's different in somewhere like Europe. But in Canada, most of the PhD students are paid for doing TAships, not primarily through grant. Average salary is 25k/year. Take 6-10k out for tuition, that's 15-19k/year. You get a student doing so many things for less pay. I guess, if your job only requires research then you can do it.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#593

Earlier quoted context omitted.

Hard to make a more substantive response when the OP’s entire comment was a one-sentence logical fallacy. I’m not cherry-picking here. > In this case, both you and the other are speculating about the near future of a thing, neither of you knows. One of us is making a much grander claim than the other: - LLMs have limitless potential for growth; because they are not capable of something today does not mean they won’t…

The post you replied to was: > We went from 2 + 7 = 11 to "solved a frontier math problem" in 3 years, yet people don't think this will improve? All that says is that the speaker thinks models will improve past where they are today. Not that it's a logical certainty (the first thing you jumped on them for), and certainly not anything about "limitless potential for growth" (which nobody even mentioned). With replies l…

> All that says is that the speaker thinks models will improve past where they are today. Not that it's a logical certainty

Exceedingly generous interpretation in my opinion. I tend to interpret rhetorical questions of that form as “it’s so obvious that I shouldn’t even have to ask it”.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#595
post #404

I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…

LLMs can generate anything by design. LLMs can't understand what they are generating so it may be true, it may be wrong, it may be novel or it may be known thing. It doesn't discern between them, just looks for the best statistical fit. The core of the issue lies in our human language and our human assumptions. We humans have implicitly assigned phrases "truly novel" and "solving unsolved math problem" a certain mean…

> It doesn't discern between them, just looks for the best statistical fit

Of course at the lowest level, LLMs are trained on next-token prediction, and on the surface, that looks like a statistics problem. But this is an incredibly reductionist viewpoint and I don't see how it makes any empirically testable predictions about their limits. LLMs 'learned' a lot of math and science in this way.

> "truly novel" and "solving unsolved math problem"

OK again if novelty lies on a continuum, where do you draw the line? And why is it correct to draw it there and not somewhere else? It seems like you are just naming exceptionally hard research problems.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#596

I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…

> 67,383 * 426,397 = 71,371,609,051 ... You need to say why it can do some novel tasks but could never do others. Model interpretability gives us the answers. The reason LLMs can (almost) do new multiplication tasks is because it saw many multiplication problems in its training data, and it was cheaper to learn the compressed/abstract multiplication strategies and encode them as circuits in the network, rather than m…

Yup, I agree with this. So based on this, where do you draw the line between what will be possible and what will not be possible?

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#597

I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…

I might as well answer my own question, because I do think there are some coherent arguments for fundamental LLM limitations:

1. LLMs are trained on human-quality data, so they will naturally learn to mimic our limitations. Their capabilities should saturate at human or maybe above-average human performance.

2. LLMs do not learn from experience. They might perform as well as most humans on certain tasks, but a human who works in a certain field/code base etc. for long enough will internalize the relevant information more deeply than an LLM.

However I'm increasingly doubtful that these arguments are actually correct. Here are some counterarguments:

1. It may be more efficient to just learn correct logical reasoning, rather than to mimic every human foible. I stopped believing this argument when LLMs got a gold metal at the Math Olympiad.

2. LLMs alone may suffer from this limitation, but RL could change the story. People may find ways to add memory. Finally, it can't be ruled out that a very large, well-trained LLM could internalize new information as deeply as a human can. Maybe this is what's happening here:

https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com...

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#598

I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…

I think "novel" is ill defined here, perhaps. LLMs do appear to be poor general reasoners[0], and it's unclear if they'll improve here. It would be unintuitive for them to be good at this, given that we know exactly how they're implemented - by looking at text and then building a statistical model to predict the next token. From this, if we wanted to commit to LLMs having generalizable knowledge, we'd have to assume…

> I think "novel" is ill defined here

That's exactly my point. When people say "LLMs will never do something novel," they seem to be leaning on some vague, ill-defined notion of novelty. The burden of proof is then to specify what degree of novelty is unattainable and why.

As for evidence that they can do novel things, there is plenty:

1. I really did ask Gemini to multiply 167,383 * 426,397 before posting this question. It answered correctly.

2. SVGs of pelicans riding bicycles

3. People use LLMs to write new apps/code every day

4. LLMs have achieved gold-medal performance on Math Olympiad problems that were not publicly available

5. LLMs have solved open problems in physics and mathematics [0,1]

That is as far as they have advanced so far. What's next? Where is the limit? All I want to say is that I don't know, and neither do you :).

[0] https://www.reddit.com/r/Physics/comments/1n77h10/and_severa... (Mark Raamsdonk is a pretty famous researcher in high-energy physics, not just some random guy)

[1] https://mathstodon.xyz/@tao/115855840223258103

[2] https://news.ycombinator.com/item?id=47497757

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#599

Earlier quoted context omitted.

No True Novelty, No True Understanding, etc. The problem with these bromides is not that they're wrong, it's that they're not even wrong. They're predictive nulls. What observable differences can we expect between an entity with True Understanding and an entity without True Understanding? It's a theological question, not a scientific one. I'm not an AI booster by any means, but I do strongly prefer we address the que…

Agreed. We should be asking what the machines measurably can or can't do. If it can't be measured, then it doesn't matter from an engineering standpoint. Does it have a soul? Can't measure it, so it doesn't matter.

That's a bit too pessimistic. Often times you can productively find some measurable proxy for the thing you care about but can't measure. Turing's test is a famous example, of that.

Sometimes you only have a one-sided proxy. Eg I can't tell you whether Claude has a soul, but I'm fairly sure my dishwasher ain't.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#600

I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…

I might as well answer my own question, because I do think there are some coherent arguments for fundamental LLM limitations: 1. LLMs are trained on human-quality data, so they will naturally learn to mimic our limitations. Their capabilities should saturate at human or maybe above-average human performance. 2. LLMs do not learn from experience. They might perform as well as most humans on certain tasks, but a human…

I studied philosophy focusing on the analytic school and proto-computer science. LLMs are going to force many people start getting a better understanding about what "Knowledge" and "Truth" are, especially the distinction between deductive and inductive knowledge.

Math is a perfect field for machine learning to thrive because theoretically, all the information ever needed is tied up in the axioms. In the empirical world, however, knowledge only moves at the speed of experimentation, which is an entirely different framework and much, much slower, even if there are some areas to catch up in previous experimental outcomes.

Having a focus in philosophy of language is something I genuinely never thought would be useful. It’s really been helpful with LLMs, but probably not in the way most people think. I’d say that folks curious should all be reading Quine, Wittgenstein’s investigations, and probably Austin.

Post reply on HN