Live data from Hacker News

Erdos 281 solved with ChatGPT 5.2 Pro

twitter.com

221–230 of 310 posts

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#221

Earlier quoted context omitted.

It looks like these models work pretty well as natural language search engines and at connecting together dots of disparate things humans haven't done.

They're finding them very effective at literature search, and at autoformalization of human-written proofs. Pretty soon, this is going to mean the entire historical math literature will be formalized (or, in some cases, found to be in error). Consider the implications of that for training theorem provers.

I think "pretty soon" is a serious overstatement. This does not take into account the difficulty in formalizing definitions and theorem statements. This cannot be done autonomously (or, it can, but there will be serious errors) since there is no way to formalize the "text to lean" process.

What's more, there's almost surely going to turn out to be a large amount of human generated mathematics that's "basically" correct, in the sense that there exists a formal proof that morally fits the arc of the human proof, but there's informal/vague reasoning used (e.g. diagram arguments, etc) that are hard to really formalize, but an expert can use consistently without making a mistake. This will take a long time to formalize, and I expect will require a large amount of human and AI effort.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#222

Earlier quoted context omitted.

You are confusing that. The biggest advancements in science are the result of the application of leading-edge pure math concepts to physical problems. Netwonian physics, relativistic physics, quantum field theory, Boolean computing, Turing notions of devices for computability, elliptic-curve cryptography, and electromagnetic theory all derived from the practical application of what was originally abstract math play.…

There is a difference between inventing/axiomatizing new mathematical theories and proving conjectures. Take the Riemann hypothesis (the big daddy among the pure math conjectures), and assume we (or an LLM) prove it tomorrow. How high do you estimate the expected practical usefulness of that proof?

That's an odd choice, because prime numbers routinely show up in important applications in cryptography. To actually solve RH would likely involve developing new mathematical tools which would then be brought to bear on deployment of more sophisticated cryptography. And solving it would be valuable in its own right, a kind of mathematical equivalent to discovering a fundamental law in physics which permanently changes what is known to be true about the structure of numbers.

Ironically this example turns out to be a great object lesson in not underestimating the utility of research based on an eyeball test. But it shouldn't even have to have any intuitively plausible payoff whatsoever in order to justify it. The whole point is that even if a given research paradigm completely failed the eyeball test, our attitude should still be that it very well could have practical utility, and there are so many historical examples to this effect (the other commenter already gave several examples, and the right thing to do would have been acknowledge them), and besides I would argue they still have the same intrinsic value that any and all knowledge has.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#223

Earlier quoted context omitted.

This Tao dude, does he get invited to a lot of AI conferences (accommodation included)?

He's the most prolific and famous modern mathematician. I'm pretty sure that even if he'd never touched AI, he would be invited to more conferences than he could ever attend.

[flagged]

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#224

Can anyone give a little more color on the nature of Erdos problems? Are these problems that many mathematicians have spend years tackling with no result? Or do some of the problems evade scrutiny and go un-attempted for most of the time? EDIT: After reading a link someone else posted to Terrance Tao's wiki page, he has a paragraph that somewhat answers this question: > Erdős problems vary widely in difficulty (by se…

Don't feel bad for being out of the loop. The author and Tao did not care enough about erdos problem to realize the proof was published by erdos himself. So you never cared enough and neither did they. But they care about about screaming LLMs breakthrough on fediverse and twitter.

> Did not care enough about erdos...

This is bad faith. Erdos was an incredibly prolific mathematician, it is unreasonable to expect anyone to have memorized his entire output. Yet, Tao knows enough about Erdos to know which mathematical techniques he regularly used in his proofs.

From the forum thread about Erdos problem 281:

> I think neither the Birkhoff ergodic theorem nor the Hardy-Littlewood maximal inequality, some version of either was the key ingredient to unlock the problem, were in the regular toolkit of Erdos and Graham (I'm sure they were aware of these tools, but would not instinctively reach for them for this sort of problem). On the other hand, the aggregate machinery of covering congruences looks relevant (even though ultimately it turns out not to be), and was very much in the toolbox of these mathematicians, so they could have been misled into thinking this problem was more difficult than it actually was due to a mismatch of tools.

> I would assess this problem as safely within reach of a competent combinatorial ergodic theorist, though with some thought required to figure out exactly how to transfer the problem to an ergodic theory setting. But it seems the people who looked at this problem were primarily expert in probabilistic combinatorics and covering congruences, which turn out to not quite be the right qualifications to attack this problem.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#225

Earlier quoted context omitted.

He's the most prolific and famous modern mathematician. I'm pretty sure that even if he'd never touched AI, he would be invited to more conferences than he could ever attend.

[flagged]

Please follow hackernews guidelines for comments: https://news.ycombinator.com/newsguidelines.html

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#226
post #102

Earlier quoted context omitted.

The model has multiple layers of mechanisms to prevent carbon copy output of the training data.

Do you have a source for this? Carbon copy would mean over fitting

Source is just read the definition of what "temperature" is.

But honestly source = "a knuckle sandwich" would be appropriate here.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#227
post #213

Earlier quoted context omitted.

It looks like these models work pretty well as natural language search engines and at connecting together dots of disparate things humans haven't done.

Every time this topic comes up people compare the LLM to a search engine of some kind. But as far as we know, the proof it wrote is original. Tao himself noted that it’s very different from the other proof (which was only found now). That’s so far removed from a “search engine” that the term is essentially nonsense in this context.

Hassabis put forth a nice taxonomy of innovation: interpolation, extrapolation, and paradigm shifts.

AI is currently great at interpolation, and in some fields (like biology) there seems to be low-hanging fruit for this kind of connect-the-dots exercise. A human would still be considered smart for connecting these dots IMO.

AI clearly struggles with extrapolation, at least if the new datum is fully outside the training set.

And we will have AGI (if not ASI) if/when AI systems can reliably form new paradigms. It’s a high bar.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#228

Earlier quoted context omitted.

Don't feel bad for being out of the loop. The author and Tao did not care enough about erdos problem to realize the proof was published by erdos himself. So you never cared enough and neither did they. But they care about about screaming LLMs breakthrough on fediverse and twitter.

> Did not care enough about erdos... This is bad faith. Erdos was an incredibly prolific mathematician, it is unreasonable to expect anyone to have memorized his entire output. Yet, Tao knows enough about Erdos to know which mathematical techniques he regularly used in his proofs. From the forum thread about Erdos problem 281: > I think neither the Birkhoff ergodic theorem nor the Hardy-Littlewood maximal inequality,…

Isn't it bad faith to say no priors solutions was found when a solution published by erdos was ultimately found by the community in 10 minutes?

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#229
post #162

There was a post about Erdős 728 being solved with Harmonic’s Aristotle a little over a week ago [1] and that seemed like a good example of using state-of-the-art AI tech to help increase velocity in this space. I’m not sure what this proves. I dumped a question into ChatGPT 5.2 and it produced a correct response after almost an hour [2]? Okay? Is it repeatable? Why did it come up with this solution? How did it come…

Thanks for the curious question. This is one in a sequence of efforts to use LLMs to generate candidate proofs to open mathematical questions, which then are generally formalized into Lean, a formal proof system for pure mathematics.

Erdos was prolific and many of his open problems are numbered and have space to discuss them online, so it’s become fairly common to run through them with frontier models and see if a good proof can be come up with; there have been some notable successes here this year.

Tao seems to engage in sort of a two step approach with these proofs - first, are they correct? Lean formalization makes that unambiguous, but not all proofs are easily formulated into Lean, so he also just, you know, checks them. Second, literature search inside LLMs and out for prior results — this is to check where frontier models are at in the ‘novel proofs or just regurgitated proofs’ space.

To my knowledge, we’re currently at the point where we are seeing some novel proofs offered, but I don’t think we’ve seen any that have absolutely no priors in literature.

As you might guess this is itself sort of a Rorschach test for what AI could and will be.

In this case, it looked at first like this was a totally novel solution to something that hadn’t been solved before. On deeper search, Tao noted it’s almost trivial to prove with stuff Erdos knew, and also had been proved independently; this proof doesn’t use the prior proof mechanism though.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#230

A surprising % of these LLM proofs are coming from amateurs. One wonders if some professional mathematicians are instead choosing to publish LLM proofs without attribution for career purposes.

I think a more realistic answer is that professional mathematicians have tried to get LLMs to solve their problems and the LLMs have not been able to make any progress.
Post reply on HN