Live data from Hacker News

“Erdos problem #728 was solved more or less autonomously by AI”

mathstodon.xyz

91–100 of 385 posts

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#91

Earlier quoted context omitted.

I agree only with the part about reconfiguring existing proofs. That's the value here. It is still likely very tedious to confirm what the LLMs say, but at least it's better than waiting for humans to do this half of the work. For all topics that can be expressed with language, the value of LLMs is shuffling things around to tease out a different perspective from the humans reading the output. This is the only realis…

> It is still likely very tedious to confirm what the LLMs say, A large amount of Tao's work is around using AI to assist in creating Lean proofs. I'm generally on the more skeptical side of things regarding LLMs and grand visions, but assisting in the creation of Lean proofs is a huge area of opportunity for LLMs and really could change mathematics in fundamental ways. One naive belief many people have is that proof…

> One naive belief many people have is that proofs should be "intelligible" but it's increasingly clear this is not the case.

That’s one of the main reason why I did not pursue an academic math career. The pure joy of solving exam problems with elegant proofs is very hard to get on harder problems.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#92

Can anyone with specific knowledge in a sophisticated/complex field such as physics or math tell me: do you regularly talk to AI models? Do feel like there's anything to learn? As a programmer, I can come to the AI with a problem and it can come up with a few different solutions, some I may have thought about, some not. Are you getting the same value in your work, in your field?

They are good for a jump start on literature search, for sure.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#93

Earlier quoted context omitted.

You're understanding correctly, this is back and forth between Aristotle and ChatGPT and a (very smart) user.

I'm not sure i understand the wild hype here in this thread then. Seems exactly like the tests at my company where even frontier models are revealed to be very expensive rubber ducks, but completely fails with non experts or anything novel or math heavy. Ie. they mirror the intellect of the user but give you big dopamine hits that'll lead you astray.

The proof is ai generated?

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#94
post #59

Earlier quoted context omitted.

Please don't respond to a bad comment by breaking the site guidelines yourself. That only makes things worse. https://news.ycombinator.com/newsguidelines.html

It genuinely seemed to me that they were looking for empirical reproductions of a formal proof, which is a nonsensical demand and objection given what formal proofs are. My question was spurred on by this and genuine. I now see in the other subthread what they mean.

It may be that there wasn't enough information in your comment for me to read its intent correctly. I thought you were taking a snarky swipe at the other commenter—especially because most people on HN can be presumed to know what a formal proof is.

If that was the case, I apologize for misreading you! If you're interested, I can point you to https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... for past explanations about this type of misunderstanding and how to avoid it in the future.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#95

Earlier quoted context omitted.

Aristotle is an LLM system.

"Aristotle integrates three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver" It is far more than an LLM, and math != "language".

> Aristotle integrates three main components (...)

The second one being backed by a model.

> It is far more than an LLM

It's an LLM with a bunch of tools around it, and a slightly different runtime that ChatGPT. It's "only" that, but people - even here, of all places - keep underestimating just how much power there is in that.

> math != "language".

How so?

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#96
post #52

[flagged]

We need you to stop posting shallow dismissals and cynical, curmudgeonly, and snarky comments. We asked you about this just recently, but it's still most of what you're posting. You're making the site worse by doing this, right at the point where it's most vulnerable these days. Your comment here is a shallow dismissal of exactly the type the HN guidelines ask users to avoid here: " Please don't post shallow dismissa…

Out of curiosity of someone who missed out on this, what is the site vulnerable to?

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#97

Can anyone with specific knowledge in a sophisticated/complex field such as physics or math tell me: do you regularly talk to AI models? Do feel like there's anything to learn? As a programmer, I can come to the AI with a problem and it can come up with a few different solutions, some I may have thought about, some not. Are you getting the same value in your work, in your field?

I talk to them (math research in algebraic geometry) not really helpful outside of literature search unfortunately. Others around me get a lot more utility so it varies. (Most powerful model i tried was Gemini 2.5 deep think and Gemini 3.0 pro) not sure if the new gpts are much better

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#98
post #52

Earlier quoted context omitted.

We need you to stop posting shallow dismissals and cynical, curmudgeonly, and snarky comments. We asked you about this just recently, but it's still most of what you're posting. You're making the site worse by doing this, right at the point where it's most vulnerable these days. Your comment here is a shallow dismissal of exactly the type the HN guidelines ask users to avoid here: " Please don't post shallow dismissa…

Out of curiosity of someone who missed out on this, what is the site vulnerable to?

Cynical, curmudgeonly, dismissive comments that ruin it as a place for curiosity.

If you're interested, https://news.ycombinator.com/item?id=46515507 and https://news.ycombinator.com/item?id=46508115 are other places I wrote about this recently.

It's the biggest problem facing HN, in my opinion.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#99

Based on Tao’s description of how the proof came about - a human is taking results backwards and forwards between two separate AI tools and using an AI tool to fill in gaps the human found? I don’t think it can really be said to have occurred autonomously then? Looks more like a 50/50 partnership with a super expert human one the one side which makes this way more vague in my opinion - and in line with my own AI test…

I had the impression Tao/community weren't even finding the gaps, since they mentioned using an automatic proof verifier. And that the main back and forth involved re-reading Erdos' paper to find out the right problem Erdos intended. So more like 90/10 LLM/human. Maybe I misread it.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#100

Earlier quoted context omitted.

You're understanding correctly, this is back and forth between Aristotle and ChatGPT and a (very smart) user.

I'm not sure i understand the wild hype here in this thread then. Seems exactly like the tests at my company where even frontier models are revealed to be very expensive rubber ducks, but completely fails with non experts or anything novel or math heavy. Ie. they mirror the intellect of the user but give you big dopamine hits that'll lead you astray.

Yes, the contributions of the people promoting the AI should be considered, as well as the people who designed the Lean libraries used in-the-loop while the AI was writing the solution. Any talk of "AGI" is, as always, ridiculous.

But speaking as a specialist in theorem proving, this result is pretty impressive! It would have likely taken me a lot longer to formalize this result even if it was in my area of specialty.

Post reply on HN