Live data from Hacker News

“Erdos problem #728 was solved more or less autonomously by AI”

mathstodon.xyz

121–130 of 385 posts

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#121

Reconfiguring existing proofs in ways that have been tedious or obscured from humans, or using well framed methods in novel ways, will be done at superhuman speeds, and it'll unlock all sorts of capabilities well before we have to be concerned about AGI. It's going to be awesome to see what mathematicians start to do with AI tools as the tools become capable of truly keeping up with what the mathematicians want from…

If this isn't AGI, what is? It seems unavoidable that an AI which can prove complex mathematical theorems would lead to something like AGI very quickly.

Tao has a comment relevant to that question:

"I doubt that anything resembling genuine "artificial general intelligence" is within reach of current #AI tools. However, I think a weaker, but still quite valuable, type of "artificial general cleverness" is becoming a reality in various ways.

By "general cleverness", I mean the ability to solve broad classes of complex problems via somewhat ad hoc means. These means may be stochastic or the result of brute force computation; they may be ungrounded or fallible; and they may be either uninterpretable, or traceable back to similar tricks found in an AI's training data. So they would not qualify as the result of any true "intelligence". And yet, they can have a non-trivial success rate at achieving an increasingly wide spectrum of tasks, particularly when coupled with stringent verification procedures to filter out incorrect or unpromising approaches, at scales beyond what individual humans could achieve.

This results in the somewhat unintuitive combination of a technology that can be very useful and impressive, while simultaneously being fundamentally unsatisfying and disappointing - somewhat akin to how one's awe at an amazingly clever magic trick can dissipate (or transform to technical respect) once one learns how the trick was performed.

But perhaps this can be resolved by the realization that while cleverness and intelligence are somewhat correlated traits for humans, they are much more decoupled for AI tools (which are often optimized for cleverness), and viewing the current generation of such tools primarily as a stochastic generator of sometimes clever - and often useful - thoughts and outputs may be a more productive perspective when trying to use them to solve difficult problems."

This comment was made on Dec. 15, so I'm not entirely confident he still holds it?

https://mathstodon.xyz/@tao/115722360006034040

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#122
I work at Harmonic, the company behind Aristotle.

To clear up a few misconceptions:

- Aristotle uses modern AI techniques heavily, including language modeling.

- Aristotle can be guided by an informal (English) proof. If the proof is correct, Aristotle has a good chance at translating it into Lean (which is a strong vote of confidence that your English proof is solid). I believe that's what happened here.

- Once a proof is formalized into Lean (assuming you have formalized the statement correctly), there is no doubt that the proof is correct. This is the core of our approach: you can do a lot of (AI-driven) search, and once you find the answer you are certain it's correct no matter how complex the solution is.

Happy to answer any questions!

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#123

Earlier quoted context omitted.

> It is still likely very tedious to confirm what the LLMs say, A large amount of Tao's work is around using AI to assist in creating Lean proofs. I'm generally on the more skeptical side of things regarding LLMs and grand visions, but assisting in the creation of Lean proofs is a huge area of opportunity for LLMs and really could change mathematics in fundamental ways. One naive belief many people have is that proof…

Math is the tip of the iceberg. If it can do proofs, it can do anything.

I don’t have proofs to solve every day, but I have to cycle the dishwasher. I eagerly await.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#124
post #83

Earlier quoted context omitted.

The goalposts are still the same. We want to be able to independently verify that an AI can do something instead of just hearing such a claim from a corporation that is absolutely willing to lie through their teeth if it gets them money.

Terrance Tao isn’t part of any AI corporation though? He’s purely a celebrated academic telling us this checks out.

[deleted]

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#125

Earlier quoted context omitted.

I'm not sure i understand the wild hype here in this thread then. Seems exactly like the tests at my company where even frontier models are revealed to be very expensive rubber ducks, but completely fails with non experts or anything novel or math heavy. Ie. they mirror the intellect of the user but give you big dopamine hits that'll lead you astray.

Yes, the contributions of the people promoting the AI should be considered, as well as the people who designed the Lean libraries used in-the-loop while the AI was writing the solution. Any talk of "AGI" is, as always, ridiculous. But speaking as a specialist in theorem proving, this result is pretty impressive! It would have likely taken me a lot longer to formalize this result even if it was in my area of specialty…

> Any talk of "AGI" is, as always, ridiculous.

How did you arrive at "ridiculous"? What we're seeing here is incredible progress over what we had a year ago. Even ARC-AGI-2 is now at over 50%. Given that this sort of process is also being applied to AI development itself, it's really not clear to me that humans would be a valuable component in knowledge work for much longer.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#126
post #94

Earlier quoted context omitted.

It genuinely seemed to me that they were looking for empirical reproductions of a formal proof, which is a nonsensical demand and objection given what formal proofs are. My question was spurred on by this and genuine. I now see in the other subthread what they mean.

It may be that there wasn't enough information in your comment for me to read its intent correctly. I thought you were taking a snarky swipe at the other commenter—especially because most people on HN can be presumed to know what a formal proof is. If that was the case, I apologize for misreading you! If you're interested, I can point you to https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... for past exp…

Thank you, I'll try to keep it in mind. I'll admit that the curtness of my original question was not just you misreading it, but it did (also) come from a place of genuine confusion.

For what it's worth, it's not even that I don't see merit to their points. I'm just unable to trust that they're being genuine, not the least for how they conduct themselves (which I only fault them for so much). This also impacts my ability to reason about their points clearly.

Sadly, I'm not able to pitch any systematic solutions.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#127

Earlier quoted context omitted.

Math is the tip of the iceberg. If it can do proofs, it can do anything.

I don’t have proofs to solve every day, but I have to cycle the dishwasher. I eagerly await.

Things like that it can’t do. But your job is more likely to be a target. Depends on what you do though.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#128
post #83

2026 should be interesting. This stuff is not magic, and progress is always going to be gradual with solutions to less interesting or "easier" problems first, but I think we're going to see more milestones like this with AI able to chip away around the edges of unsolved mathematics. Of course, that will require a lot of human expertise too: even this one was only "solved more or less autonomously by AI (after some fe…

The goalposts are still the same. We want to be able to independently verify that an AI can do something instead of just hearing such a claim from a corporation that is absolutely willing to lie through their teeth if it gets them money.

Not disagreeing with you, but I don't think Tao is blowing this out of proportion either. I think it's a pretty reasonable way of saying, "Hey, AI is now capable of something it wasn't able to do before".

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#129

I work at Harmonic, the company behind Aristotle. To clear up a few misconceptions: - Aristotle uses modern AI techniques heavily, including language modeling. - Aristotle can be guided by an informal (English) proof. If the proof is correct, Aristotle has a good chance at translating it into Lean (which is a strong vote of confidence that your English proof is solid). I believe that's what happened here. - Once a pr…

First congrats!

Sometimes when I'm using new LLMs I'm not sure if it’s a step forward or just benchmark hacking, but formalized math results always show that the progress is real and huge.

When do you think Harmonic will reach formalizing most (even hard) human written math?

I saw an interview with Christian Szegedy (your competitor I guess) that he believes it will be this year.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#130

It took Andrew Wiles 7 years of intense work to solve Fermat's Last Theorem. The METR institute predicts that the length of tasks AI agents can complete doubles every 7 months. We should expect it to take until 2033 before AI solves Clay Institute-level problems with 50% reliability.

There is an ongoing effort to formalize a modern, streamlined proof of FLT in Lean, with all the needed prereqs. It's estimated that it will take approx. 5 years, but perhaps AI will lead to some meaningful speedup.

What I'm hoping to see is high volume automated formalization of the math literature, with the goal of formalizing (or finding flaws in) the entire thing.

And once we have that formalized corpus, it's all set up as training data for moving forward.

Post reply on HN