Live data from Hacker News

Erdos 281 solved with ChatGPT 5.2 Pro

twitter.com

121–130 of 310 posts

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#121
post #29
post #11

Earlier quoted context omitted.

My take is that a huge part of human intelligence is pattern matching. We just didn’t understand how much multidimensional geometry influenced our matches

Yes, it could be that intelligence is essentially a sophisticated form of recursive, brute force pattern matching. I'm beginning to think the Bitter Lesson applies to organic intelligence as well, because basic pattern matching can be implemented relatively simply using very basic mathematical operations like multiply and accumulate, and so it can scale with massive parallelization of relatively simple building block…

Intelligence is almost certainly a fundamentally recursive process.

The ability to think about your own thinking over and over as deeply as needed is where all the magic happens. Counterfactual reasoning occurs every time you pop a mental stack frame. By augmenting our stack with external tools (paper, computers, etc.), we can extend this process as far as it needs to go.

LLMs start to look a lot more capable when you put them into recursive loops with feedback from the environment. A trillion tokens worth of "what if..." can be expended without touching a single token in the caller's context. This can happen at every level as many times as needed if we're using proper recursive machinery. The theoretical scaling around this is extremely favorable.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#122
post #10

FWIW, I just gave Deepseek the same prompt and it solved it too (much faster than the 41m of ChatGPT). I then gave both proofs to Opus and it confirmed their equivalence. The answer is yes. Assume, for the sake of contradiction, that there exists an \(\epsilon > 0\) such that for every \(k\), there exists a choice of congruence classes \(a_1^{(k)}, \dots, a_k^{(k)}\) for which the set of integers not covered by the f…

> I then gave both proofs to Opus and it confirmed their equivalence. You could have just rubber-stamped it yourself, for all the mathematical rigor it holds. The devil is in the details, and the smallest problem unravels the whole proof.

How dare you question the rigor of the venerable LLM peer review process! These are some of the most esteemed LLMs we are talking about here.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#123

Earlier quoted context omitted.

> It wasn't AI generated. You're lying: https://www.pangram.com/history/94678f26-4898-496f-9559-8c4c... Not that I needed pangram to tell me that, it's obvious slop.

I must be a bot because I love existential dread, that's a great phrase. I feel like they trigger a lot on literate prose.

Sad times when the only remaining way to convince LLM luddites of somebody’s humanity is bad writing.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#124
post #11
post #9

This is crazy. It's clear that these models don't have human intelligence, but it's undeniable at this point that they have _some_ form of intelligence.

My take is that a huge part of human intelligence is pattern matching. We just didn’t understand how much multidimensional geometry influenced our matches

Intelligence is hallucination that happens to produce useful results in the real world.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#125

A surprising % of these LLM proofs are coming from amateurs. One wonders if some professional mathematicians are instead choosing to publish LLM proofs without attribution for career purposes.

It's probably from the perennial observation

"This LLM is kinda dumb in the thing I'm an expert in"

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#126
Personally, I'd prefer if the AI models would start with a proof of their own statements. Time and again, SOTA frontier models told me: "Now you have 100% correct code ready for production in enterprise quality." Then I run it and it crashes. Or maybe the AI is just being tongue-in-cheek?

Point in case: I just wanted to give z.ai a try and buy some credits. I used Firefox with uBlock and the payment didn't go through. I tried again with Chrome and no adblock, but now there is an error: "Payment Failed: p.confirmCardPayment is not a function." The irony is, that this is certainly vibe-coded with z.ai which tries to sell me how good they are but then not being able to conclude the sale.

And we will get lots more of this in the future. LLMs are a fantastic new technology, but even more fantastically over-hyped.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#127
post #10

FWIW, I just gave Deepseek the same prompt and it solved it too (much faster than the 41m of ChatGPT). I then gave both proofs to Opus and it confirmed their equivalence. The answer is yes. Assume, for the sake of contradiction, that there exists an \(\epsilon > 0\) such that for every \(k\), there exists a choice of congruence classes \(a_1^{(k)}, \dots, a_k^{(k)}\) for which the set of integers not covered by the f…

"Since \(U_{k+1} \subseteq U_k\), the sets \(U_k\) are decreasing and periodic, and their intersection \(U = \bigcap_{k \ge 1} U_k\) has density \(d = \lim_{k \to \infty} d_k \ge \epsilon\)."

Is this enough? Let $U_k$ be the set of integers such that their remainder mod 6^n is greater or equal to 2^n for all 1<n<k. Density of each $U_k$ is more than 1/2 I think but not the intersection (empty) right?

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#128
post #126

Personally, I'd prefer if the AI models would start with a proof of their own statements. Time and again, SOTA frontier models told me: "Now you have 100% correct code ready for production in enterprise quality." Then I run it and it crashes. Or maybe the AI is just being tongue-in-cheek? Point in case: I just wanted to give z.ai a try and buy some credits. I used Firefox with uBlock and the payment didn't go through…

You get AIs to prove their code is correct in precisely the same ways you get humans to prove their code is correct. You make them demonstrate it through tests or evidence (screenshots, logs of successful runs).

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#129
post #93
post #9

This is crazy. It's clear that these models don't have human intelligence, but it's undeniable at this point that they have _some_ form of intelligence.

Well, Alpha Go and Stockfish can beat you at their games. Why shouldn't these models beat us at math proofs?

Alpha go and stockfish were specifically designed and trained to win at those games.

Re: Erdos 281 solved with ChatGPT 5.2 Pro

#130

Out of curiosity why has the LLM math solving community been focused on the Erdos problems over other open problems? Are they of a certain nature where we would expect LLMs to be especially good at solving them?

People like checking items off of lists.
Post reply on HN