Live data from Hacker News

Epoch confirms GPT5.4 Pro solved a frontier math open problem

epoch.ai

571–580 of 744 posts

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#571

I feel like this single image perfectly sums up the entire thread here: https://trapatsas.eu/sites/llm-predictions/

Yes, and no matter when "now" is, the doubters will always see in their mind's eye the flat line extending to the right.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#572
post #229

Earlier quoted context omitted.

> programming languages / tools better suited for LLM strengths The bitter lesson is that the best languages / tools are the ones for which the most quality training data exists, and that's pretty much necessarily the same languages / tools most commonly used by humans. > Correct code not nice looking code "Nice looking" is subjective, but simple, clear, readable code is just as important as ever for projects to be l…

>> simple, clear, readable code is just as important as ever for projects to be long-term successful Is it though? I'm a long-time code purist, but I am beginning to wonder about the assumptions underlying our vocation.

I guess it's hard to tell until we see more long-term AI-generated project, but many of the ones we have so far (OpenClaw and OpenCode for instance) are well-known for their stability issues, and it seems "even more AI" is not about to fix that.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#574
post #380
post #282

Earlier quoted context omitted.

That may be, and we can debate the level of novelty, but it is novel, because this exact proof didn't exist before, something which many claim was not possible with AI. In fact, just a few years ago, based on some dabbling in NLP a decade ago, I myself would not have believed any of this was remotely possible within the next 3 - 5 decades at least. I'm curious though, how many novel Math proofs are not close enough t…

Well, for one the proof would have to use actual proof techniques. What really happened here was that the LLM produced a python script that generated examples of hypergraphs that served as proof by example . And the only thing that has been verified are these examples. The LLM also produced a lot of mathematical text that has not been analyzed.

I see, thanks for the explanation!

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#575

Earlier quoted context omitted.

Yes! I call these the "it's just a stochastic parrot" crowd. Ironically, they are the stochastic parrots, because they're confidently repeating something that they read somehwere and haven't examined critically.

That would not be stochastic, just parroting

It's not deterministic therefore it's stochastic.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#576

Earlier quoted context omitted.

OK, its not a logical fallacy, its a false assumption. The belief in the inevitability of progress is a bad assumption. Especially if you assume a particular technology will keep advancing.

We won't know if his assumption is false until time passes and moves future speculation into the empirical present.

A possibility is not a fact. Assuming a possibility will happen is not justified. Therefore it is false as an assumption, even if it is true it is a possiblity.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#577
post #454

Earlier quoted context omitted.

"Truly novel" is fast becoming a True Scotsman.

No True Novelty, No True Understanding, etc. The problem with these bromides is not that they're wrong, it's that they're not even wrong. They're predictive nulls. What observable differences can we expect between an entity with True Understanding and an entity without True Understanding? It's a theological question, not a scientific one. I'm not an AI booster by any means, but I do strongly prefer we address the que…

Well said. That's exactly what has been rubbing me the wrong way with all those "LLMs can never *really* think, ya know" people. Once we pass some level of AI capability (which we perhaps already did?), it essentially turns into an unfalsifiable statement of faith.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#578

Earlier quoted context omitted.

It's only because humans came up with a problem, worked with the ai and verified the result that this achievement means anything at all. An ai "checking its own work" is practically irrelevant when they all seem to go back and forth on whether you need the car at the carwash to wash the car. Undoubtedly people have been passing this set of problems to ai's for months or years and have gotten back either incorrect res…

Funding a few PhDs for a year costs orders of magnitude more than it did to solve this problem in inference costs. Also, this has been active research for some time. Or I guess the people working on it are just not as good as a random bunch of students? It's amazing the lengths that people go to maintain their worldview, even if it means belittling hardworking people. I take it you're not a mathematician. This is an…

> Funding a few PhDs for a year costs orders of magnitude more than it did to solve this problem in inference costs.

I don't think PhD students are sitting around and solving one problem for a year. Also PhD students are way cheaper

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#579
post #13

I like to imagine that the number of consumed tokens before a solution is found is a proxy for how difficult a problem is, and it looks like Opus 4.6 consumed around 250k tokens. That means that a tricky React refactor I did earlier today at work was about half as hard as an open problem in mathematics! :)

I don't think so. I went through the output of Opus 4.6 vs GPT 5.4 pro. Both are given different directions/prompts. Opus 4.6 was asked to test and verify many things. Opus 4.6 tried in many different ways and the chain of thoughts are more interesting to me.

Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem

#580

Earlier quoted context omitted.

Inference costs are heavily subsidised. My point was that we've spent trillions collectively on ai, and so far we have a few new proofs. It's been active research but the problem estimates only 5-10 people are even aware that it is a problem. I wrote "math phd's" not "random students", but regardless, I wouldn't know how you interpreted my statement that people could have discovered without ai this as "belittling the…

>> we've spent trillions Source? This sounds like hyperbole. The entire US GDP is low tens of trillions.

From various online estimates, i would estimate global ai spend just since 2020 at $2T. Some projections estimate that we might spend that per year starting next year. To the extent that many of these projects will be cancelled or shelved, capital is beginning to take stock of the feasibility of clawing back even the original investments. openai is apparently doubling its staff, but whether these are sales or (prompt?) engineering jobs, the biggest hypemongers are themselves unable to reduce headcount even with unlimited "at-cost" ai inference.

Comparing total ai spend to the value added of producing a few new maths/sciences proofs is unfair since ai is doing more than maths proofs, but for comparison one can estimate the total spent to date on mathematicians and associated costs (buildings, experiments etc). I would very roughly estimate that the total cost of all mathematics to date since 1600 is less than what we've spent on ai to date, and the results from investment in mathematicians are incomparable to a few derivative extensions of well-established ideas. For less than a few trillion we have all of mathematics. For an additional 2T dollars, we have trivial advancements that no one really cares about.

Post reply on HN