Live data from Hacker News

Epoch confirms GPT5.4 Pro solved a frontier math open problem

epoch.ai

491–500 of 744 posts

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#491

Earlier quoted context omitted.

Linear regression has well characterized mathematical properties. But we don't know the computational limits of stacked transformers. And so declaring what LLMs can't do is wildly premature.

> And so declaring what LLMs can't do is wildly premature. The opposite is true as well. Emergent complexity isn’t limitless. Just like early physicists tried to explain the emergent complexity of the universe through experimentation and theory, so should we try to explain the emergent complexity of LLMs through experimentation and theory. Specifically not pseudoscience, though.

>so should we try to explain the emergent complexity of LLMs through experimentation and theory.

Physicists had the real world to verify theories and explanations against.

So far anyone 'explaining the emergent complexity of LLMs through experimentation and theory' is essentially just making stuff up nobody can verify.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#492
post #454

Earlier quoted context omitted.

If LLMs can come up with formerly truly novel solutions to things, and you have a verification loop to ensure that they are actual proper solutions, I don't understand why you think they could never come up with solutions to impressive problems, especially considering the thread we are literally on right now? That seems like a pure assertion at this point that they will always be limited to coming up with truly novel…

"Truly novel" is fast becoming a True Scotsman.

No True Novelty, No True Understanding, etc.

The problem with these bromides is not that they're wrong, it's that they're not even wrong. They're predictive nulls.

What observable differences can we expect between an entity with True Understanding and an entity without True Understanding? It's a theological question, not a scientific one.

I'm not an AI booster by any means, but I do strongly prefer we address the question of AI agent intelligence scientifically rather than theologically.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#493
post #404

I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…

LLMs can generate anything by design. LLMs can't understand what they are generating so it may be true, it may be wrong, it may be novel or it may be known thing. It doesn't discern between them, just looks for the best statistical fit. The core of the issue lies in our human language and our human assumptions. We humans have implicitly assigned phrases "truly novel" and "solving unsolved math problem" a certain mean…

> LLMs can't understand what they are generating

You don't understand what "understanding" means. I'm sure you can't explain it. You are probably just hallucinating the feeling of understanding it.

> Some of us at least, think that truly novel means something truly novel and important, something significant. Like, I don't know...

Yeah.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#494

Earlier quoted context omitted.

Sure, that's true as well. But I don't see this as a substantive response given that the only people making unsupported claims in this thread are those trying to deflate LLM capabilities.

So, to review this thread - OP asked for someone to make a logical argument for the separation of “training” from “model” - I made the argument - You cherry picked an argument against my specific example and made an appeal to emergent complexity - I pointed out that emergent complexity isn’t limitless - “the only people making unsupported claims in this thread are those trying to deflate LLM capabilities”

You made a pretty nonsensical argument, pretty much seems like the big standard for these arguments.

What does linear regression have to do with the limitations of a stacked transfer ? Absolutely nothing. This is the problem here. You don't know shit and just make up whatever. You can see people doing the same thing in GPT-1, 2, 3, 4 threads all telling us why LLMs will never be able to do thing it manages to do later.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#496

Earlier quoted context omitted.

what bothers me is not that this issue will certainly disappear now that it has been identified, but that that we have yet to identify the category of these "stupid" bugs ...

We already know exactly what causes these bugs. They are not a fundamental problem of LLMs, they are a problem of tokenizers. The actual model simply doesn't get to see the same text that you see. It can only infer this stuff from related info it was trained on. It's as if someone asked you how many 1s there are in the binary representation of this text. You'd also need to convert it first to think it through, or use…

> It's as if someone asked you how many 1s there are in the binary representation of this text.

I'm actually kinda pleased with how close I guessed! I estimated 4 set bits per character, which with 491 characters in your post (including spaces) comes to 1964.

Then I ran your message through a program to get the actual number, and turns out it has 1800 exactly.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#497

Earlier quoted context omitted.

> I don't see this getting better. We went from 2 + 7 = 11 to "solved a frontier math problem" in 3 years, yet people don't think this will improve?

I’ve seen this style of take so much that I’m dying for someone to name a logical fallacy for it, like “appeal to progress” or something. Step away from LLMs for a second and recognize that “Yesterday it was X, so today it must be X+1” is such a naive take and obviously something that humans so easily fall into a trap of believing (see: flying cars).

Logical fallacies are vastly overrated. Unless the conversation is formal logic in the first place, "logical fallacies" are just a way to apply quick pattern matching to dismiss people without spending time on more substantive responses. In this case, both you and the other are speculating about the near future of a thing, neither of you knows.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#498

Earlier quoted context omitted.

All of the data is still in the prompt, you are just asking the model to do a simple transform. I think there are examples of what you’re looking for, but this isn’t one.

> All of the data is still in the prompt, you are just asking the model to do a simple transform. LLMs can use data in their prompt. They can also use data in their context window. They can even augment their context with persisted data. You can also roll out LLM agents, each one with their role and persona, and offload specialized tasks with their own prompts, context windows, and persisted data, and even tools to g…

I was in no way dismissing it -- I was refuting the above claim that they "generate things they have not seen before"

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#499

Earlier quoted context omitted.

All of the data is still in the prompt, you are just asking the model to do a simple transform. I think there are examples of what you’re looking for, but this isn’t one.

I agree that this isn't a very interesting example, but your statement is: "just asking the model to do a simple transform". If you assert that it understand when you ask it things like that, how could anything it produces not fall under the "already in the model" umbrella?

I didn't say it wasn't an interesting example -- i said it wasn't an example of LLMs generating things they have not seen before.

> how could anything it produces not fall under the "already in the model" umbrella

It doesn't. That is the point of my comment.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#500
post #473

Earlier quoted context omitted.

If LLMs can come up with formerly truly novel solutions to things, and you have a verification loop to ensure that they are actual proper solutions, I don't understand why you think they could never come up with solutions to impressive problems, especially considering the thread we are literally on right now? That seems like a pure assertion at this point that they will always be limited to coming up with truly novel…

It probably can, but won't realize that and it won't be efficient in that. LLM can shuffle tokens for an enormous number of tries and eventually come up with something super impressive, though as you yourself have mentioned, we would need to have a mandatory verification loop, to filter slop from good output and how to do it outside of some limited areas is a big question. But assuming we have these verification loop…

> We never had a big demand to define how humans are intelligent or conscious etc, since it is too hard and was relegated to a some frontier researchers. And with LLMs we now do have such demand but the science wasn't ready. So we are all collectively searching in the dark, trying to define if we are different from these programs if not how. I certainly can't do that. I do know that LLMs are useful, but I also suspect that AI (aka AGI nowadays) is not yet reached.

Alternative perspective: the science may not have been ready, so instead we brute-forced the problem, through training of LLMs. Consider what the overall goal function of LLM training is: it's predicting tokens that continue given input in a way that makes sense to humans - in fully general meaning of this statement.

It's a single training process that gives LLMs the ability to parse plain language - even if riddled with 1337-5p34k, typos, grammar errors, or mixing languages - and extract information from it, or act on it; it's the same single process that makes it equally good at writing code and poetry, at finding bugs in programs, inconsistencies in data, corruptions in images, possibly all at once. It's what makes LLMs good at lying and spotting lies, even if input is a tree of numbers.

(It's also why "hallucinations" and "prompt injection" are not bugs, but fundamental facets of what makes LLMs useful. They cannot and will not be "fixed", any more than you can "fix" humans to be immune to confabulation and manipulation. It's just the nature of fully general sytems.)

All of that, and more, is encoded in this simple goal function: if a human looks at the output, will they say it's okay or nonsense? We just took that and thrown a ton of compute at it.

Post reply on HN