Live data from Hacker News

Epoch confirms GPT5.4 Pro solved a frontier math open problem

epoch.ai

561–570 of 744 posts

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#561

Earlier quoted context omitted.

> The “good deal of evidence” is everywhere. The proof is in the pudding. I'm open! Please, by all means. > the blog article (not an actual paper?) rightfully derides benchmarks and then…creates a benchmark? The blog article is a review of benchmarking methodologies and the issues involved by a PhD neuroscientist who works directly on large language models and their applications to neuroscience and cognition, it's pr…

> The “good deal of evidence” is everywhere. The proof is in the pudding. I'm open! Please, by all means. sure here are but a few: [1] you get smooth gains in reasoning with more RL train-time compute and more test-time compute (o1) [2] DeepSeek-R1 showed that RL on verifiable rewards produces behavior like backtracking, adaptation, reflection, etc. [3] SWE-Bench is a relatively decent benchmark and perf here is cont…

Thanks, this is great and I'll have quite a bit to read here.

> people like to move goalposts whenever a new result comes out, which is silly. Could AI systems do this 2 years ago? No. I don’t know how people don’t look at robust trends in performance improvement, combined with verifiable RL rewards, and can’t understand where things are going.

I don't think it's goal post moving to acknowledge improvements but still reject the conclusion that AI has reached a specific milestone if those improvements don't justify the position. I doubt anyone sensible is rejecting improvements.

> But you pretend as though this is definitive evidence that “LLMs are poor general reasoners”.

I don't think I've ever made any definitive claims at all, quite the contrary - I've tried to express exactly how open I am to what you're saying. As I've said, I'm a functionalist, and I already am largely supportive of reductive intelligence, so I'm exactly the type of person who would be sympathetic to what you're saying.

> "That’s how science is done" is a bit of an oversimplification

Of course, but I don't think it's too much to ask for to have a theory and evidence. I don't need a lined up series of papers that all start with perfectly syllogisms and then map to well controlled RCTs or whatever. Just an "I think this accounts for it, here's how I support that".

> The claim that there is no framework + no real tests is just not true anymore.

I didn't say it wasn't true, to be clear, I asked for it. Again, I'm sympathetic to the view at a glance so I simply need a way to reason about it.

No need for a complete view, I'd never expect such a thing.

> The model is: reasoning is not inherently human, it’s mathematical.

Well, hand wringing perhaps, but I'd say it's maybe mathematical, computational, structural, functional, whatever - I think we're on the same page here regardless.

> It falls easily within the purview of RL, statistics, representation, optimization, etc, and to claim otherwise would require evidence.

Sure, but I grant that, in fact I believe it entirely. But that doesn't mean that every mathematical construct exhibits the function of intelligence.

> What is the robust model for reasoning in humans again? Simulations and models — what are these? Interventative analysis — we can’t do this with LLMs? Falsifying test cases — what would satisfy you here beyond everything I’ve presented above?

Sorry, I'm not fully understanding this framing. We can do those things with LLMs, and it's hard to say what I would be satisfied. In general, I'd be satisfied with a theory that (a) accounts for the data (b) has supporting evidence (c) does not contradict any major prior commitments. I don't think (c) will be an issue here.

> You say “brains are intelligent” ==> “intelligence is an emergent property of cells zapping” is absurd,

Because intelligence could have been a property of our brains being wet, or roundish, or it could have been a property of our spines, or maybe some force we hadn't discovered, or a soul, etc. We formed a theory, it accounted for observations, we performed tests, we've modeled things, etc, and so the theories we've adopted have been extremely successful and I think hold up quite well. But certainly we didn't go "the brain has electricity, the brain is intelligent, therefor electricity in the brain is what drives intelligence".

> Brains _are_ made up of real, physical atoms organized into molecules organized into cells organized into a coordinated system, and…that’s it? What’s missing here?

Certainly nothing on my world view.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#562

Earlier quoted context omitted.

Math and coding competition problems are easier to train because of strict rules and cheap verification. But once you go beyond that to less defined things such as code quality, where even humans have hard time putting down concrete axioms, they start to hallucinate more and become less useful. We are missing the value function that allowed AlphaGo to go from mid range player trained on human moves to superhuman by p…

LLMs can often guess the final answer, but the intermediate proof steps are always total bunk. When doing math you only ever care about the proof, not the answer itself.

Not in this case: the LLM wrote the entire paper, and anyway the proof was the answer.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#563
post #539

Earlier quoted context omitted.

Ok, I'll bite. Show me an LLM that comes up with a new math operator. Or which will come up with theory of relativity if only Newton physics is in its training dataset. That it could remix existing ideas which leads to novel insights is expected, however the current LLMs can't come up with paradigm shifts that require novel insights. Even humans have a rather limited time they can come up with novel insights (when th…

How many humans have been born until now and how many Einsteins have been born? And in how many hundreds of thousands of years?

The point is that humans do have some edge compared to current LLMs which are essentially next token predictors. If we all start relying on current AI and stop thinking, we would only be able to "exhaust the remix space" of existing ideas but won't be able to do any paradigm jumps. Moreover, it's quite likely that current training sets are self-contradictory, containing Dutch books, carrying some innate error in them.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#564

Earlier quoted context omitted.

I'm following this mini-thread with interest but I've arrived here and I confess, I don't really know what your argument is. I think this all stems from you objecting to this statement: "I don't know why I am still perpetually shocked that the default assumption is that humans are somehow unique." I think you're being uncharitable in how you interpret that. Human's are unique in the most literal reading of this sente…

They're shocked that people believe that humans are unique. I explained why that shouldn't be shocking. I think I was pretty charitable here, I gave an alternative option for what they could mean in my very first reply: > Unless you mean "fundamentally unique" in some way that would persist - like "nothing could ever do what humans do". > I don't really know what your argument is. I just said that I think that we hav…

I still think you're being far too literal, which doesn't make for an interesting conversation.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#565

Earlier quoted context omitted.

They're shocked that people believe that humans are unique. I explained why that shouldn't be shocking. I think I was pretty charitable here, I gave an alternative option for what they could mean in my very first reply: > Unless you mean "fundamentally unique" in some way that would persist - like "nothing could ever do what humans do". > I don't really know what your argument is. I just said that I think that we hav…

I still think you're being far too literal, which doesn't make for an interesting conversation.

I'm open to hearing how you think I should be interpreting things. I don't really think I'm being too literal, it certainly hasn't been the case that they've suggested my interpretation is wrong, and I've provided two interpretations (one that I totally grant).

What's the better interpretation of their position?

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#566

Earlier quoted context omitted.

Someone actually mathed out infinite monkeys at infinite typewriters, and it turns out, it is a great example of how misleading probabilities are when dealing with infinity: "Even if every proton in the observable universe (which is estimated at roughly 1080) were a monkey with a typewriter, typing from the Big Bang until the end of the universe (when protons might no longer exist), they would still need a far greate…

> So no. LLMs are not brute force dummies. We are seeing increasingly emergent behavior in frontier models. Woah! That was a leap. "We are seeing ... emergent behaviors" does not follow from "it's not brute force". It is unsurprising that an LLM performs better than random! That's the whole point. It does not imply emergence.

> It is unsurprising that an LLM performs better than random! That's the whole point. It does not imply emergence.

By definition, it is emergent behavior when it exhibits the ability to synthesize solutions to problems that it wasn't trained on. I.e. it can handle generalization.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#567

Earlier quoted context omitted.

> So no. LLMs are not brute force dummies. We are seeing increasingly emergent behavior in frontier models. Woah! That was a leap. "We are seeing ... emergent behaviors" does not follow from "it's not brute force". It is unsurprising that an LLM performs better than random! That's the whole point. It does not imply emergence.

> It is unsurprising that an LLM performs better than random! That's the whole point. It does not imply emergence. By definition, it is emergent behavior when it exhibits the ability to synthesize solutions to problems that it wasn't trained on. I.e. it can handle generalization.

Emergent behavior would imply that some other function was being reduced to token prediction. Behaving "better than random" ie: not just brute forcing would not qualify - token prediction is not brute forcing and we expect it to do better, it's trained to do so.

If you want to demonstrate an emergent behavior you're going to need to show that.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#568
post #481

Earlier quoted context omitted.

We have different senses of humor.

Just tell one funny thing an LLM said...

Yesterday it was "LLM's can't count R's in 'strawberry'." Today it's "LLM's can't tell jokes". Tomorrow it might be "LLM's can't do (X)", all while LLMs get better and better at every objection/challenge posed.

The problem as I see it is that you have a fundamental objection to categorizing the way LLMs do their work as in any way related to "real gosh-darn human thinking". Which I think is wrong. At the root, we are just information-processing meat that happens to have had millions of years to optimize for speed, pattern recognition, feedback, etc.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#569
post #251
post #217

Earlier quoted context omitted.

1) this is a proof by example 2) the proof is conducted by writing a python program constructing hypergraphs 3) the consensus was this was low-hanging fruit ready to be picked, and tactics for this problem were available to the LLM So really this is no different from generating any python program. There are also many examples of combinatoric construction in python training sets. It's still a nice result, but it's not…

One of the possible outcomes of this journey is that “LLMs can never do X”. Another is that X is easier than we thought.

Or that some quixotic problems nobody cared about to the extent to actually work on them do have some solution.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#570

Earlier quoted context omitted.

Funding a few PhDs for a year costs orders of magnitude more than it did to solve this problem in inference costs. Also, this has been active research for some time. Or I guess the people working on it are just not as good as a random bunch of students? It's amazing the lengths that people go to maintain their worldview, even if it means belittling hardworking people. I take it you're not a mathematician. This is an…

Inference costs are heavily subsidised. My point was that we've spent trillions collectively on ai, and so far we have a few new proofs. It's been active research but the problem estimates only 5-10 people are even aware that it is a problem. I wrote "math phd's" not "random students", but regardless, I wouldn't know how you interpreted my statement that people could have discovered without ai this as "belittling the…

>> we've spent trillions

Source? This sounds like hyperbole. The entire US GDP is low tens of trillions.

Post reply on HN