Live data from Hacker News

Erdos 281 solved with ChatGPT 5.2 Pro

twitter.com

131–140 of 310 posts

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#131

> no prior solutions found. This is no longer true, a prior solution has just been found[1], so the LLM proof has been moved to the Section 2 of Terence Tao's wiki[2]. [1] - https://www.erdosproblems.com/forum/thread/281#post-3325 [2] - https://github.com/teorth/erdosproblems/wiki/AI-contribution...

This illustrates how unimportant this problem is. A prior solution did exist, but apparently nobody knew because people didn't really care about it. If progress can be had by simply searching for old solutions in the literature, then that's good evidence the supposed progress is imaginary. And this is not the first time this has happened with an Erdős problem. A lot of pure mathematics seems to consist in solving nea…

There is still enormous value in cleaning up the long tail of somewhat important stuff. One of the great benefits of Claude Code to me is that smaller issues no longer rot in backlogs, but can be at least attempted immediately.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#133

> no prior solutions found. This is no longer true, a prior solution has just been found[1], so the LLM proof has been moved to the Section 2 of Terence Tao's wiki[2]. [1] - https://www.erdosproblems.com/forum/thread/281#post-3325 [2] - https://github.com/teorth/erdosproblems/wiki/AI-contribution...

[flagged]

[deleted]

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#134

A surprising % of these LLM proofs are coming from amateurs. One wonders if some professional mathematicians are instead choosing to publish LLM proofs without attribution for career purposes.

I'm actually not sure what the right attribution method would be. I'd lean towards single line on acknowledgements? Because you can use it for example @ every lemma during brainstorming but it's unclear the right convention is to thank it at every lemma...

Anecdotally, I, as a math postdoc, think that GPT 5.2 is much stronger qualitatively than anything else I've used. Its rate of hallucinations is low enough that I don't feel like the default assumption of any solution is that it is trying to hide a mistake somewhere. Compared with Gemini 3 whose failure mode when it can't solve something is always to pretend it has a solution by "lying"/ omitting steps/making up theorems etc... GPT 5.2 usually fails gracefully and when it makes a mistake it more often than not can admit it when pointed out.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#135

Out of curiosity why has the LLM math solving community been focused on the Erdos problems over other open problems? Are they of a certain nature where we would expect LLMs to be especially good at solving them?

I guess they are at a difficulty where it's not too hard (unlike millennium prize problems), is fairly tightly scoped (unlike open ended research), and has some gravitas (so it's not some obscure theorem that's only unproven because of it's lack of noteworthiness).

I actually don't think the reason is that they are easier than other open math problems. I think it's more that they are "elementary" in the sense that the problems usually don't require a huge amount of domain knowledge to state.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#136

Can anyone give a little more color on the nature of Erdos problems? Are these problems that many mathematicians have spend years tackling with no result? Or do some of the problems evade scrutiny and go un-attempted for most of the time? EDIT: After reading a link someone else posted to Terrance Tao's wiki page, he has a paragraph that somewhat answers this question: > Erdős problems vary widely in difficulty (by se…

Erdos was an incredibly prolific mathematician, and one of his quirks is that he liked to collect open problems and state new open problems as a challenge to the field. Many of the problems he attached bounties to, from $5 to $10,000.

The problems are a pretty good metric for AI, because the easiest ones at least meet the bar of "a top mathematician didn't know how to solve this off the top of his head" and the hardest ones are major open problems. As AI progresses, we will see it slowly climb the difficulty ladder.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#137
post #93

Earlier quoted context omitted.

Well, Alpha Go and Stockfish can beat you at their games. Why shouldn't these models beat us at math proofs?

Alpha go and stockfish were specifically designed and trained to win at those games.

And we can train models specifically at math proofs? I think only difference is that math is bigger....

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#138
post #10

FWIW, I just gave Deepseek the same prompt and it solved it too (much faster than the 41m of ChatGPT). I then gave both proofs to Opus and it confirmed their equivalence. The answer is yes. Assume, for the sake of contradiction, that there exists an \(\epsilon > 0\) such that for every \(k\), there exists a choice of congruence classes \(a_1^{(k)}, \dots, a_k^{(k)}\) for which the set of integers not covered by the f…

Opus isn't a good choice for anything math-related; it's worse at math than the latest ChatGPT and Gemini Pro.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#139
post #102

Earlier quoted context omitted.

The model has multiple layers of mechanisms to prevent carbon copy output of the training data.

forgive the skepticism, but this translates directly to "we asked the model pretty please not to do it in the system prompt"

It's mind boggling if you think about the fact they're essential "just" statistical models

It really contextualizes the old wisdom of Pythagoras that everything can be represented as numbers / math is the ultimate truth

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#140

Earlier quoted context omitted.

Why not plan for a future where a lot of non-trivial tasks are automated instead of living on the edge with all this anxiety?

[flagged]

I mean.. LLMs have hit a pretty hard wall a while ago, with the only solution being throwing monstrous compute at eking out the remaining few percent improvement (real world, not benchmarks). That's not to mention hallucinations / false paths being a foundational problem.

LLMs will continue to get slightly better in the next few years, but mainly a lot more efficient. Which will also mean better and better local models. And grounding might get better, but that just means less wrong answers, not better right answers.

So no need for doomerism. The people saying LLMs are a few years away from eating the world are either in on the con or unaware.

Post reply on HN