Live data from Hacker News

Epoch confirms GPT5.4 Pro solved a frontier math open problem

epoch.ai

541–550 of 744 posts

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#541

Earlier quoted context omitted.

The things it didn't, which you then helpfully spelled out.

Do you believe my brief overview of the problem will help Claude identify the specific undocumented functions required for my solution? Is that how you think data gets fed back into models during training?

Yes. I don't think you appreciate just how much information your comments provide. You just told us (and Claude) what the interesting problems are, and confirmed both the existence of relevant undocumented functions, and that they are the right solution to those problems. What you didn't flag as interesting, and possible challenges you did not mention (such as these APIs being flaky, or restricted to Apple first-party use, or such) is even more telling.

Most hard problems are hard because of huge uncertainty around what's possible and how to get there. It's true for LLMs as much as it is for humans (and for the same reasons). Here, you gave solid answers to both, all but spelling out the solution.

ETA:

> Is that how you think data gets fed back into models during training?

No, one comment chain on a niche site is not enough.

It is, however, how the data gets fed into prompt, whether by user or autonomously (e.g. RAG).

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#542
post #536

Earlier quoted context omitted.

How can people look at - clear generalizability - insane growth rates (go back and look at where we were maybe 2 years ago and then consider the already signed compute infrastructure deals coming online) And still say with a straight face that this is some kind of parlor trick or monkeys with typewriters. we don’t need to run LLMs for years. The point is look at where we are today and consider performance gets 10x ch…

I was talking about highest difficulty problems only, in the scope of that comment. Sure at mundane tasks they are useful and we optimizing that constantly. But for super hard tasks, there is no situation when you just dump a few papers for context add a prompt and LLM will spit out correct answer. It's likely that a lead on such project would need to additionally train LLM on their local dataset, then parse through…

This is just not true. Maybe it will be true if you increase the problem difficulty in concert with model performance? You don't need fine tuning for this and you haven't for years now. Reasoning performance for now may be SOMEWHAT brittle but again look at where we have come from in like 2 years. Then also consider the logical next steps

- better context compression (already happening) + memory solutions that extend the effective context length [memory _is_ compression]

- continual learning systems (likely already prototyped)

- these domains are _verifiable_ which I think just seems to confuse people. RL in verifiable domains takes you farther and farther. Training data is a bootstrap to get to a starting point, because RL from scratch is too inefficient.

agents can already deal with large codebases and datasets, just like any SWE, DS or researcher.

and yes! If you throw more compute at a problem you will get better solutions! But you are missing the point: for the frontier solutions, which changes with every model update, you of course need to eek out as much performance as you can, which requires a large amount of test time compute. But what you can do _without_ this is continually improving. The pattern _already in place_ is that at first you need an extreme amount of compute, then the next model iterations need far less compute to reach that same solution, etc etc. The costs + compute requirements to perform a particular task decrease exponentially.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#543

Earlier quoted context omitted.

> I don't know why I am still perpetually shocked that the default assumption is that humans are somehow unique. Because, empirically, we have numerous unique and differentiable qualities, obviously. Plenty of time goes into understanding this, we have a young but rigorous field of neuroscience and cognitive science. Unless you mean "fundamentally unique" in some way that would persist - like "nothing could ever do w…

>But that does not mean that I have to think that the appearance of intelligence always is intelligence, or that an LLM/ Agent is doing what humans do. You can think whatever you want, but an untestable distinction is an imaginary one.

First of all, that's not true. Not every position has to be empirically justified. I can reason about a position in all sorts of ways without testing. Here's an obvious example that requires no test at all:

1. Functional properties seem to arise from structural properties

2. Brains and LLMs have radically different structural properties

3. Two constructs with radically, fundamentally different structural properties are less likely to have identical functional properties

Therefor, my confidence in the belief that brains and LLMs should have identical functional properties is lowered by some amount, perhaps even just ever so slightly.

Not something I feel like fleshing out or defending, just an example of how I could reason about a position without testing it.

Second, I never said it wasn't testable.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#544
post #539

I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…

Ok, I'll bite. Show me an LLM that comes up with a new math operator. Or which will come up with theory of relativity if only Newton physics is in its training dataset. That it could remix existing ideas which leads to novel insights is expected, however the current LLMs can't come up with paradigm shifts that require novel insights. Even humans have a rather limited time they can come up with novel insights (when th…

How many humans have been born until now and how many Einsteins have been born? And in how many hundreds of thousands of years?

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#545

Earlier quoted context omitted.

Yeah. What else would it be ? A brain capable of doing that was clearly the result of evolutionary pressures.

But there is no evolutionary pressure for the Poincaré conjecture, we were never optimized for that in particular, unlike these kinds of LLMs.

Of course it is evolution. What else could it be?

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#547

I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…

The hardest part about any creativity is hiding your influences

This is poetry.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#548

Earlier quoted context omitted.

Hard to make a more substantive response when the OP’s entire comment was a one-sentence logical fallacy. I’m not cherry-picking here. > In this case, both you and the other are speculating about the near future of a thing, neither of you knows. One of us is making a much grander claim than the other: - LLMs have limitless potential for growth; because they are not capable of something today does not mean they won’t…

The post you replied to was: > We went from 2 + 7 = 11 to "solved a frontier math problem" in 3 years, yet people don't think this will improve? All that says is that the speaker thinks models will improve past where they are today. Not that it's a logical certainty (the first thing you jumped on them for), and certainly not anything about "limitless potential for growth" (which nobody even mentioned). With replies l…

Better put than I could have.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#549

Earlier quoted context omitted.

Logical fallacies are vastly overrated. Unless the conversation is formal logic in the first place, "logical fallacies" are just a way to apply quick pattern matching to dismiss people without spending time on more substantive responses. In this case, both you and the other are speculating about the near future of a thing, neither of you knows.

OK, its not a logical fallacy, its a false assumption. The belief in the inevitability of progress is a bad assumption. Especially if you assume a particular technology will keep advancing.

We won't know if his assumption is false until time passes and moves future speculation into the empirical present.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#550

Earlier quoted context omitted.

> I don't see this getting better. We went from 2 + 7 = 11 to "solved a frontier math problem" in 3 years, yet people don't think this will improve?

I’ve seen this style of take so much that I’m dying for someone to name a logical fallacy for it, like “appeal to progress” or something. Step away from LLMs for a second and recognize that “Yesterday it was X, so today it must be X+1” is such a naive take and obviously something that humans so easily fall into a trap of believing (see: flying cars).

The comment doesn't say it must be X+1. It implies it will improve which I would say is a pretty safe bet.
Post reply on HN