How are academics going to assess AI-coauthored research for appointment and promotion?
“Erdos problem #728 was solved more or less autonomously by AI”
101–110 of 385 posts
Re: “Erdos problem #728 was solved more or less autonomously by AI”
#102Earlier quoted context omitted.
I'm not sure i understand the wild hype here in this thread then. Seems exactly like the tests at my company where even frontier models are revealed to be very expensive rubber ducks, but completely fails with non experts or anything novel or math heavy. Ie. they mirror the intellect of the user but give you big dopamine hits that'll lead you astray.
The proof is ai generated?
"Aristotle integrates three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver"
Not saying it's not an amazing setup, i just don't understand the word "AI" being used like this when it's the setup / system that's brilliant in conjunction with absolute experts.
Re: “Erdos problem #728 was solved more or less autonomously by AI”
#103It took Andrew Wiles 7 years of intense work to solve Fermat's Last Theorem. The METR institute predicts that the length of tasks AI agents can complete doubles every 7 months. We should expect it to take until 2033 before AI solves Clay Institute-level problems with 50% reliability.
Re: “Erdos problem #728 was solved more or less autonomously by AI”
#104This almost implies mathematicians aren’t some ungodly geniuses if something as absolutely dumb as an LLM can solve these problems via blind pattern matching. Meanwhile I can’t get Claude code to fix its own shit to save my life.
Re: “Erdos problem #728 was solved more or less autonomously by AI”
#105Earlier quoted context omitted.
> beyond LLMs to specialized approached Do you mean that in this case, it was not a LLM?
It could not be done without Aristotle ( https://arxiv.org/pdf/2510.01346 ), as clearly described in Tao's posts.
Re: “Erdos problem #728 was solved more or less autonomously by AI”
#106Earlier quoted context omitted.
"Aristotle integrates three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver" It is far more than an LLM, and math != "language".
> Aristotle integrates three main components (...) The second one being backed by a model. > It is far more than an LLM It's an LLM with a bunch of tools around it, and a slightly different runtime that ChatGPT. It's "only" that, but people - even here, of all places - keep underestimating just how much power there is in that. > math != "language". How so?
LLM = Large Language Model. Large refers to both the number of parameters (and in practice, depth) of the model, and also implicitly the amount of data used for training, and "language" means human (i.e. written, spoken) language. A Vision Transformer is not an LLM because it is trained on images, and AlphaFold is not an LLM because it is trained molecular configurations.
Aristotle works heavily with formalized LEAN statements and expressions. While you can certainly argue this is a language of sorts, it is not at all the same "language" as the "language" in LLMs. Calling Aristotle an "LLM" just because it has a transformer is more misleading than truthful, because every other single aspect of it is far more clever and involved.
Re: “Erdos problem #728 was solved more or less autonomously by AI”
#1072026 should be interesting. This stuff is not magic, and progress is always going to be gradual with solutions to less interesting or "easier" problems first, but I think we're going to see more milestones like this with AI able to chip away around the edges of unsolved mathematics. Of course, that will require a lot of human expertise too: even this one was only "solved more or less autonomously by AI (after some fe…
Re: “Erdos problem #728 was solved more or less autonomously by AI”
#1082026 should be interesting. This stuff is not magic, and progress is always going to be gradual with solutions to less interesting or "easier" problems first, but I think we're going to see more milestones like this with AI able to chip away around the edges of unsolved mathematics. Of course, that will require a lot of human expertise too: even this one was only "solved more or less autonomously by AI (after some fe…
The goalposts are still the same. We want to be able to independently verify that an AI can do something instead of just hearing such a claim from a corporation that is absolutely willing to lie through their teeth if it gets them money.
Re: “Erdos problem #728 was solved more or less autonomously by AI”
#109Earlier quoted context omitted.
The proof is ai generated?
Eh? The text reads: "Aristotle integrates three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver" Not saying it's not an amazing setup, i just don't understand the word "AI" being used like this when it's the setup / system that's brilliant in conjunction with absolute experts.
Re: “Erdos problem #728 was solved more or less autonomously by AI”
#110Earlier quoted context omitted.
Out of curiosity of someone who missed out on this, what is the site vulnerable to?
Cynical, curmudgeonly, dismissive comments that ruin it as a place for curiosity. If you're interested, https://news.ycombinator.com/item?id=46515507 and https://news.ycombinator.com/item?id=46508115 are other places I wrote about this recently. It's the biggest problem facing HN, in my opinion.
Surprisingly the latest increase in polarization around generative AI has impacted Hacker News the least our of all tech social spaces.