Live data from Hacker News

Epoch confirms GPT5.4 Pro solved a frontier math open problem

epoch.ai

461–470 of 744 posts

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#461
post #255

Earlier quoted context omitted.

Unfair - human beats AI in this comparison, as human will instantly answer "I don't know" instead of yelling a random number. Or at best "I don't know, but maybe I can find out" and proceed to finding out/ But he is unlikely to shout "6" because he heard this number once when someone talked about light.

> human will instantly answer "I don't know" instead of yelling a random number. Seems that you never worked with Accenture consultants?

Fair.

Yet this can be filtered with fixed rules, like "output produced by corporate structures is untrusted random data".

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#462
post #332

Earlier quoted context omitted.

It is like not trusting someone who attained highest score in some exam by by-hearting the whole text book, to do the corresponding job. Not very hard to understand.

Yet we do that all the time by hiring based on GPA/degree.

Do you hire or screen based on them?

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#463

Earlier quoted context omitted.

When it comes to LLMs doing novel things, is it just the infinite monkey theorem[0] playing out at an accelerated rate, helped along by the key presses not being truly random? Surely if we tell the LLM to do enough stuff, something will look novel, but how much confirmation bias is at play? Tens of millions of people are using AI and the biggest complaint is hallucinations. From the LLMs perspective, is there any dif…

This argument doesn't go the way you want it to go. Billions of people exist, but maybe a few tens of thousands produce novel knowledge. That's a much worse rate than LLMs.

I’m not sure how we equate the number of humans to AI to determine a success rate.

We also can’t ignore than it was humans who thought up this problem to give to the AI. Thinking has two parts, asking and answering questions. The AI needed the human to formulate and ask the question to start. AI isn’t just dropping random discoveries on us that we haven’t even thought of, at least not that I’ve seen.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#464
post #379

Earlier quoted context omitted.

So....what is your point?

Generations have grown and died in the time since your concern was first expressed. The world continues. Culture adapts.

Did I say it will be the end of the world?

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#465

Earlier quoted context omitted.

it is not the assumption that humans are unique. it is that statistical models cannot really think out of the box most of the time

And you know that humans aren't statistical models how?

because they would be more logical

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#467
post #404

I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…

LLMs can generate anything by design. LLMs can't understand what they are generating so it may be true, it may be wrong, it may be novel or it may be known thing. It doesn't discern between them, just looks for the best statistical fit. The core of the issue lies in our human language and our human assumptions. We humans have implicitly assigned phrases "truly novel" and "solving unsolved math problem" a certain mean…

> It doesn't discern between them, just looks for the best statistical fit.

Why this is not true for humans?

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#468

Earlier quoted context omitted.

Math and coding competition problems are easier to train because of strict rules and cheap verification. But once you go beyond that to less defined things such as code quality, where even humans have hard time putting down concrete axioms, they start to hallucinate more and become less useful. We are missing the value function that allowed AlphaGo to go from mid range player trained on human moves to superhuman by p…

> I don't see this getting better. We went from 2 + 7 = 11 to "solved a frontier math problem" in 3 years, yet people don't think this will improve?

Self driving

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#469

Earlier quoted context omitted.

You could have just checked the math yourself, you know.

My pocket calculator says the same thing and it doesn't even have training data.

Huh, casually flexing with a sentient pocket calculator!!!

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#470
post #404

Earlier quoted context omitted.

LLMs can generate anything by design. LLMs can't understand what they are generating so it may be true, it may be wrong, it may be novel or it may be known thing. It doesn't discern between them, just looks for the best statistical fit. The core of the issue lies in our human language and our human assumptions. We humans have implicitly assigned phrases "truly novel" and "solving unsolved math problem" a certain mean…

> It doesn't discern between them, just looks for the best statistical fit. Why this is not true for humans?

We can't tell yet if that is true, partially true, or false for humans. We do know that LLM can't do anything else besides that (I mean as a fundamental operating principle).
Post reply on HN