Live data from Hacker News

“Erdos problem #728 was solved more or less autonomously by AI”

mathstodon.xyz

251–260 of 385 posts

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#251
post #248

When Deep Blue beat Kaspaorov, it was not the end of career for human players. But since mathematics is not a sport with human players, what are the career prospects for mathematicians or mathematics-like fields?

I think its worth saying two things:

1. This result is very far from showing something like "human mathematicians are no longer needed to advance mathematics".

2. Even if it did show that, as long as we need humans trained in understanding maths, since "professional mathematicians" are mostly educators, they probably aren't going anywhere.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#252

Earlier quoted context omitted.

How do you verify that the AI translation to Lean is a correct formalization of the problem? In other fields, generative AI is very good at making up plausible sounding lies, so I'm wondering how likely that is for this usage.

That's what's covered by the "assuming you have formalized the statement correctly" parenthetical. Given a formal statement of what you want, Lean can validate that the steps in a (tedious) machine-readable purported proof are valid and imply the result from accepted axioms. This is not AI, but a tiny, well reviewed kernel that only accepts correct formal logic arguments. So, if you have a formal statement that you'v…

Soo, it can definitively tell you that 42 is correct Answer to the Ultimate Question of Life, The Universe, and Everything. It just can't tell you if you're asking the right question.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#254

Earlier quoted context omitted.

I'm not sure i understand the wild hype here in this thread then. Seems exactly like the tests at my company where even frontier models are revealed to be very expensive rubber ducks, but completely fails with non experts or anything novel or math heavy. Ie. they mirror the intellect of the user but give you big dopamine hits that'll lead you astray.

This accurately mirrors my experience. It never - so far - has happened that the AI brought any novel insight at the level that I would see as an original idea. Presumably the case of TFA is different but the normal interaction is that that the solution to whatever you are trying to solve is a millimeter away from your understanding and the AI won't bridge that gap until you do it yourself and then it will usually pr…

That problem is not clearly stated, so if you’re pasting that into an AI verbatim you won’t get the answer you’re looking for.

My guess is: first move the weights to the middle, and only then remove them.

However “weights” and “bar” might confuse both machines and people into thinking that this is related to weight lifting, where there’s two stops on the bar preventing the weights from being moved to the middle.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#255

Earlier quoted context omitted.

They probably need to be able to read and understand the lean language.

I can read and understand e.g. Python, but I have seen subtle bugs that were hard to spot in code generated by AI. At least the last time I tried coding agents (mid 2025), it was often easier to write the code myself then play "spot the bug" with whatever was generated. I don't know anything about Lean, so I was wondering if there were similar pitfalls here.

If you want to check the statement, you only have to read the type. The proof itself you don’t have to read at all

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#256
post #154

Earlier quoted context omitted.

You're looking for the practical answer, but philosophically it isn't possible to translate an informal statement into a formal one 'correctly'. It is informal, ie, vaguely specified. The only certain questions are if the formal axioms and results are interesting which is independent of the informal formalisation and that can only be established by inspecting the the proof independently of the informal spec.

Philosophically, this is not true in general , but that's for trivial reasons: "how many integers greater than 7 are blue?" doesn't correspond to a formal question. It is absolutely true in many specific cases. Most problems posed by a mathematician will correspond to exactly one formal proposition, within the context of a given formal system. This problem is unusual, in that it was originally misspecified.

I suppose there's no formally defined procedure that accepts a natural language statement and outputs either its formalization or "misspecified". And "absolutely true" means "the vast majority of mathematicians agree that there's only one formal proposition that corresponds to this statement".

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#257

Earlier quoted context omitted.

That's what's covered by the "assuming you have formalized the statement correctly" parenthetical. Given a formal statement of what you want, Lean can validate that the steps in a (tedious) machine-readable purported proof are valid and imply the result from accepted axioms. This is not AI, but a tiny, well reviewed kernel that only accepts correct formal logic arguments. So, if you have a formal statement that you'v…

Soo, it can definitively tell you that 42 is correct Answer to the Ultimate Question of Life, The Universe, and Everything. It just can't tell you if you're asking the right question.

No, it can tell you that 42 is the answer to (some lean statement), but not what question that lean statement encodes.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#258
post #83

Earlier quoted context omitted.

The goalposts are still the same. We want to be able to independently verify that an AI can do something instead of just hearing such a claim from a corporation that is absolutely willing to lie through their teeth if it gets them money.

Terrance Tao isn’t part of any AI corporation though? He’s purely a celebrated academic telling us this checks out.

From his posts, it’s unclear who actually did the experiment. He seems to only be commenting on the results? Or am I missing something?

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#259
post #136

Earlier quoted context omitted.

In a very specialized setup, in tandem with a verifier. Just because a specialized human placed in an F-16 can fly at Mach 2.0, doesn't mean humans in general can fly.

An apt analogy. A human is a general intelligence that can fly with an F-16. What happens when we put an artificial general intelligence in an F-16? That's what happened here with this proof.

Not really. A completely unintelligent autopilot can fly an F-16. You cannot assume general intelligence from scaffolded tool-using success in a single narrow area.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#260
post #248

When Deep Blue beat Kaspaorov, it was not the end of career for human players. But since mathematics is not a sport with human players, what are the career prospects for mathematicians or mathematics-like fields?

I think its worth saying two things: 1. This result is very far from showing something like "human mathematicians are no longer needed to advance mathematics". 2. Even if it did show that, as long as we need humans trained in understanding maths, since "professional mathematicians" are mostly educators, they probably aren't going anywhere.

> ... are mostly educators, they probably aren't going anywhere

Educator business survived so far, only because they provided in-person interactive knowledge transfer and credentials - both were not possible by static sources of knowledge such as libraries and internet. But now all that is possible without involvement of human teachers.

Post reply on HN