Earlier quoted context omitted.
I am kind of joking, but I actually don't know where the flaw in my logic is. It's like one of those math proofs that 1 + 1 = 3. If I were to hazard a guess, I think that tokens spent thinking through hard math problems probably correspond to harder human thought than tokens spend thinking through React issues. I mean, LLMs have to expend hundreds of tokens to count the number of r's in strawberry. You can't tell me…
You can spend countless "tokens" solving minesweeper or sudoku. This doesn't mean that you solved difficult problems: just that the solutions are very long and, while each step requires reasoning, the difficulty of that reasoning is capped.
Epoch confirms GPT5.4 Pro solved a frontier math open problem
611–620 of 744 posts
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#612Earlier quoted context omitted.
Do you believe my brief overview of the problem will help Claude identify the specific undocumented functions required for my solution? Is that how you think data gets fed back into models during training?
Yes. I don't think you appreciate just how much information your comments provide. You just told us (and Claude) what the interesting problems are, and confirmed both the existence of relevant undocumented functions, and that they are the right solution to those problems. What you didn't flag as interesting, and possible challenges you did not mention (such as these APIs being flaky, or restricted to Apple first-part…
Lol... no. You don't know how I solved the problem and you just read everything that Claude did.
Absolutely nothing in the key part of my solution uses a single public API (and there are thousands). And you think that Claude can just "figure that out" when my HK comments gets fed back in during training?
I sincerely wish we'd see less /r/technology ridiculousness on HN.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#613Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#614Earlier quoted context omitted.
> I think "novel" is ill defined here That's exactly my point. When people say "LLMs will never do something novel," they seem to be leaning on some vague, ill-defined notion of novelty. The burden of proof is then to specify what degree of novelty is unattainable and why. As for evidence that they can do novel things, there is plenty: 1. I really did ask Gemini to multiply 167,383 * 426,397 before posting this quest…
Actually here's an even better list of progress on a number of open math problems, with plenty of caveats and exposition: https://github.com/teorth/erdosproblems/wiki/AI-contribution...
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#615Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#616I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…
Especially the lemmas:
- any statement about AI which uses the word "never" to preclude some feature from future realization is false.
- contemporary implementations have almost always already been improved upon, but are unevenly distributed.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#617I am kind of amazed at how many commenters respond to this result by confidently asserting that LLMs will never generate 'truly novel' ideas or problem solutions. > AI is a remixer; it remixes all known ideas together. It won't come up with new ideas > it's not because the model is figuring out something new > LLMs will NEVER be able to do that, because it doesn't exist It's not enough to say 'it will never be able t…
Ximm's Law applies ITT: every critique of AI assumes to some degree that contemporary implementations will not, or cannot, be improved upon. Especially the lemmas: - any statement about AI which uses the word "never" to preclude some feature from future realization is false. - contemporary implementations have almost always already been improved upon, but are unevenly distributed.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#618Earlier quoted context omitted.
LLMs can generate anything by design. LLMs can't understand what they are generating so it may be true, it may be wrong, it may be novel or it may be known thing. It doesn't discern between them, just looks for the best statistical fit. The core of the issue lies in our human language and our human assumptions. We humans have implicitly assigned phrases "truly novel" and "solving unsolved math problem" a certain mean…
> It doesn't discern between them, just looks for the best statistical fit Of course at the lowest level, LLMs are trained on next-token prediction, and on the surface, that looks like a statistics problem. But this is an incredibly reductionist viewpoint and I don't see how it makes any empirically testable predictions about their limits. LLMs 'learned' a lot of math and science in this way. > "truly novel" and "sol…
Did they? Or is it begging the question?
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#619Earlier quoted context omitted.
No True Novelty, No True Understanding, etc. The problem with these bromides is not that they're wrong, it's that they're not even wrong. They're predictive nulls. What observable differences can we expect between an entity with True Understanding and an entity without True Understanding? It's a theological question, not a scientific one. I'm not an AI booster by any means, but I do strongly prefer we address the que…
We've tested this in the small with AI art. When people believe they're viewing human-made art which is later revealed to be AI art, they feel disappointed. The actual content is incidental, the story that supports it is more important than the thing itself. It's the same mechanism behind artisanal food, artist struggles, and luxury goods. It is the metaphysical properties we attach to objects or the frames we use to…
The actual content of a work of art is the expression of lived experience. Not its form.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#620Earlier quoted context omitted.
Math and coding competition problems are easier to train because of strict rules and cheap verification. But once you go beyond that to less defined things such as code quality, where even humans have hard time putting down concrete axioms, they start to hallucinate more and become less useful. We are missing the value function that allowed AlphaGo to go from mid range player trained on human moves to superhuman by p…
> I don't see this getting better. We went from 2 + 7 = 11 to "solved a frontier math problem" in 3 years, yet people don't think this will improve?
This is disingenuous... I don't think people were impressed by GPT 3.5 because it was bad at math.
It's like saying: "We went from being unable to take off and the crew dying in a fire to a moon landing in 2 years, imagine how soon we'll have people on Mars"