Live data from Hacker News

“Erdos problem #728 was solved more or less autonomously by AI”

mathstodon.xyz

101–110 of 385 posts

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#102

Earlier quoted context omitted.

I'm not sure i understand the wild hype here in this thread then. Seems exactly like the tests at my company where even frontier models are revealed to be very expensive rubber ducks, but completely fails with non experts or anything novel or math heavy. Ie. they mirror the intellect of the user but give you big dopamine hits that'll lead you astray.

The proof is ai generated?

Eh? The text reads:

"Aristotle integrates three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver"

Not saying it's not an amazing setup, i just don't understand the word "AI" being used like this when it's the setup / system that's brilliant in conjunction with absolute experts.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#103

It took Andrew Wiles 7 years of intense work to solve Fermat's Last Theorem. The METR institute predicts that the length of tasks AI agents can complete doubles every 7 months. We should expect it to take until 2033 before AI solves Clay Institute-level problems with 50% reliability.

If you have a sufficiently strong verifier 1/100000 reliability is already enough

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#104

This almost implies mathematicians aren’t some ungodly geniuses if something as absolutely dumb as an LLM can solve these problems via blind pattern matching. Meanwhile I can’t get Claude code to fix its own shit to save my life.

You're right we're not

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#105
post #7

Earlier quoted context omitted.

> beyond LLMs to specialized approached Do you mean that in this case, it was not a LLM?

It could not be done without Aristotle ( https://arxiv.org/pdf/2510.01346 ), as clearly described in Tao's posts.

Never mind what Aristotle is, verifier llm models are definitely strong enough to verify proofs of elementary methods used here.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#106

Earlier quoted context omitted.

"Aristotle integrates three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver" It is far more than an LLM, and math != "language".

> Aristotle integrates three main components (...) The second one being backed by a model. > It is far more than an LLM It's an LLM with a bunch of tools around it, and a slightly different runtime that ChatGPT. It's "only" that, but people - even here, of all places - keep underestimating just how much power there is in that. > math != "language". How so?

Transformer != LLM. See my edited top-level post. Just because Aristotle uses a transformer doesn't mean it is an LLM, just as Vision Transformers and AlphaFold use transformers but are not LLMs.

LLM = Large Language Model. Large refers to both the number of parameters (and in practice, depth) of the model, and also implicitly the amount of data used for training, and "language" means human (i.e. written, spoken) language. A Vision Transformer is not an LLM because it is trained on images, and AlphaFold is not an LLM because it is trained molecular configurations.

Aristotle works heavily with formalized LEAN statements and expressions. While you can certainly argue this is a language of sorts, it is not at all the same "language" as the "language" in LLMs. Calling Aristotle an "LLM" just because it has a transformer is more misleading than truthful, because every other single aspect of it is far more clever and involved.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#107

2026 should be interesting. This stuff is not magic, and progress is always going to be gradual with solutions to less interesting or "easier" problems first, but I think we're going to see more milestones like this with AI able to chip away around the edges of unsolved mathematics. Of course, that will require a lot of human expertise too: even this one was only "solved more or less autonomously by AI (after some fe…

I think 2026 should see insane progress in AI for math (if not in AI generally)

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#108
post #83

2026 should be interesting. This stuff is not magic, and progress is always going to be gradual with solutions to less interesting or "easier" problems first, but I think we're going to see more milestones like this with AI able to chip away around the edges of unsolved mathematics. Of course, that will require a lot of human expertise too: even this one was only "solved more or less autonomously by AI (after some fe…

The goalposts are still the same. We want to be able to independently verify that an AI can do something instead of just hearing such a claim from a corporation that is absolutely willing to lie through their teeth if it gets them money.

Terrance Tao isn’t part of any AI corporation though? He’s purely a celebrated academic telling us this checks out.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#109

Earlier quoted context omitted.

The proof is ai generated?

Eh? The text reads: "Aristotle integrates three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver" Not saying it's not an amazing setup, i just don't understand the word "AI" being used like this when it's the setup / system that's brilliant in conjunction with absolute experts.

[deleted]

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#110
post #98

Earlier quoted context omitted.

Out of curiosity of someone who missed out on this, what is the site vulnerable to?

Cynical, curmudgeonly, dismissive comments that ruin it as a place for curiosity. If you're interested, https://news.ycombinator.com/item?id=46515507 and https://news.ycombinator.com/item?id=46508115 are other places I wrote about this recently. It's the biggest problem facing HN, in my opinion.

> It's the biggest problem facing HN, in my opinion.

Surprisingly the latest increase in polarization around generative AI has impacted Hacker News the least our of all tech social spaces.

Post reply on HN