Live data from Hacker News

“Erdos problem #728 was solved more or less autonomously by AI”

mathstodon.xyz

291–300 of 385 posts

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#291
post #24

Can anyone with specific knowledge in a sophisticated/complex field such as physics or math tell me: do you regularly talk to AI models? Do feel like there's anything to learn? As a programmer, I can come to the AI with a problem and it can come up with a few different solutions, some I may have thought about, some not. Are you getting the same value in your work, in your field?

Context: I finished a PhD in pure math in 2025 and have transitioned to being a data scientist and I do ML/stats research on the side now. For me, deep research tools have been essential for getting caught up with a quick lit review about research ideas I have now that I'm transitioning fields. They have also been quite helpful with some routine math that I'm not as familiar with but is relatively established (like s…

Do you use LLM models? Or something else?

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#292

Earlier quoted context omitted.

That problem is not clearly stated, so if you’re pasting that into an AI verbatim you won’t get the answer you’re looking for. My guess is: first move the weights to the middle, and only then remove them. However “weights” and “bar” might confuse both machines and people into thinking that this is related to weight lifting, where there’s two stops on the bar preventing the weights from being moved to the middle.

The problem is stated clearly enough that humans that we ask the question of will sooner or later see that there is an optimum and that that optimum relies on understanding. And no, the problem is not 'not clearly stated'. It is complete as it is and you are wrong about your guess. And if machines and people think this is related to weight lifting then they're free to ask follow up questions. But even in the weight l…

Illusion of transparency. You are imagining yourself asking this question, while standing in the gym and looking at the bar (or something like this). I, for example, have no idea how the weights are attached and which removal actions are allowed.

Yeah, LLMs have a tendency to run with some interpretation of a question without asking follow-up questions. Probably, it's a consequence of RLHFing them in that way.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#293
post #268

Earlier quoted context omitted.

> apparently When someone takes the time to explain undergrad-level concepts in a comment, responding with "are you an expert?" is a level of skepticism that's bordering on hostile. The person you're responding to is correct, it's rare that the theorem statement itself is particularly hard to formalize. Whatever you read likely refers to the difficulty of formalizing a proof.

> it's rare that the theorem statement itself is particularly hard to formalize That's very dependent on the problem area. For example there's a gap between high school explanation of central limit theorem and actual formalization of it. And when dealing with turing machines sometimes you'll say that something grows e.g. Omega(n), but what happens is that there's some subsequence of inputs for which it does. Generall…

Yes, if the theorem statement itself is "hard to formalize" even given our current tools, formal foundations etc. for this task, this suggests that the underlying math itself is still half-baked in some sense, and could be improved to better capture the concepts we're interested in. Much of analysis-heavy math is in that boat at present, compared to algebra.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#294

Earlier quoted context omitted.

They probably need to be able to read and understand the lean language.

I can read and understand e.g. Python, but I have seen subtle bugs that were hard to spot in code generated by AI. At least the last time I tried coding agents (mid 2025), it was often easier to write the code myself then play "spot the bug" with whatever was generated. I don't know anything about Lean, so I was wondering if there were similar pitfalls here.

As I understand it Lean is not a general purpose programming language, it is a DSL focused on formal logic verification. Bugs in a DSL are generally easier to identify and fix.

It seems one side of this argument desperately needs AI to have failed, and the other side is just saying that it probably worked but it is not as important as presented, that it is actually just a very cool working methodology going forward.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#295

Earlier quoted context omitted.

> Any talk of "AGI" is, as always, ridiculous. How did you arrive at "ridiculous"? What we're seeing here is incredible progress over what we had a year ago. Even ARC-AGI-2 is now at over 50%. Given that this sort of process is also being applied to AI development itself, it's really not clear to me that humans would be a valuable component in knowledge work for much longer.

It requires constant feedback, critical evaluation, and checks. This is not AGI, its cognitive augmentation. One that is collective, one that will accelerate human abilities far beyond what the academic establishment is currently capable of, but that is still fundamentally organic. I don't see a problem with this--AGI advocates treat machine intelligence like some sort of God that will smite non-believers and reward…

>AGI advocates treat machine intelligence like some sort of God that will smite non-believers and reward the faithful.

>The real world is not composed of rewards and punishments.

Most "AGI advocates" say that AGI is coming, sooner rather than later, and it will fundamentally reshape our world. On its own that's purely descriptive. In my experience, most of the alleged "smiting" comes from the skeptics simply being wrong about this. Rarely there's talk of explicit rewards and punishments.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#296

Earlier quoted context omitted.

The problem is stated clearly enough that humans that we ask the question of will sooner or later see that there is an optimum and that that optimum relies on understanding. And no, the problem is not 'not clearly stated'. It is complete as it is and you are wrong about your guess. And if machines and people think this is related to weight lifting then they're free to ask follow up questions. But even in the weight l…

Illusion of transparency. You are imagining yourself asking this question, while standing in the gym and looking at the bar (or something like this). I, for example, have no idea how the weights are attached and which removal actions are allowed. Yeah, LLMs have a tendency to run with some interpretation of a question without asking follow-up questions. Probably, it's a consequence of RLHFing them in that way.

And none of those details matter to solve the problem correctly. I'm purposefully not putting any answers here because I want to see if future generations of these tools suddenly see the non-obvious solution. But you are right about the fact that the details matter, one detail is mentioned very explicitly that holds the key.

If you do solve it don't post the answer.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#297

Earlier quoted context omitted.

AGI in its standard definition requires matching or surpassing humans on all cognitive tasks, not just in some, especially some where only handful of humans took a stab on.

Surely AGI would be matching humans on most tasks. To me, surpassing humans on all cognitive tasks sounds like superintelligence, while AGI "only" need to perform most, but not necessarily all, cognitive tasks at the level of a human highly capable at that task.

Super intelligence is defined as outmatching the best humans in a field, but again, on all cognitive tasks, not just a subset.

AI can already beat humans in pretty much any game like Go or Chess or many videogames, but that doesn't make it general.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#298

Earlier quoted context omitted.

> Any talk of "AGI" is, as always, ridiculous. How did you arrive at "ridiculous"? What we're seeing here is incredible progress over what we had a year ago. Even ARC-AGI-2 is now at over 50%. Given that this sort of process is also being applied to AI development itself, it's really not clear to me that humans would be a valuable component in knowledge work for much longer.

> it's really not clear to me that humans would be a valuable component in knowledge work for much longer. To me, this sounds like when we first went to the moon, and people were sure we'd be on Mars be the end of the 80's. > Even ARC-AGI-2 is now at over 50%. Any measure of "are we close to AGI" is as scientifically meaningful as "are we close to a warp drive" because all anyone has to go on at this point is pure sp…

> To me, this sounds like when we first went to the moon, and people were sure we'd be on Mars be the end of the 80's.

Unlike space colonisation, there are immediate economic rewards from producing even modest improvements in AI models. As such, we should expect much faster progress in AI than space colonisation.

But it could still turn out the same way, for all we know. I just think that's unlikely.

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#299

Earlier quoted context omitted.

Then what sort of math problem would be a milestone for you where an AI was doing something novel? Or are you just saying that solving novel problems involves remixing ideas? Well, that's true for human problem solving too.

> Then what sort of math problem would be a milestone for you where an AI was doing something novel? What? If we're discussing novel synthesis, and it's being contrasted with answer-from-search / answer-from-remix.. the problem does not matter. Only the answer and the originality of the approach. Connecting two fields that were not previously connected is novel, or applying a new kind of technique to an old problem.…

If you squint hard enough, every new thing is an example of "answer-from-search / answer-from-remix". Solving any Erdős problem in this manner was largely seen as unthinkable just a year ago.

>the problem does not matter.

Really? All of the other Erdős problems? Millennium Problems? Anything at all? This gets us directly into the territory of "nothing can convince us otherwise".

Re: “Erdos problem #728 was solved more or less autonomously by AI”

#300

Earlier quoted context omitted.

Let me put it like this: I expect AI to replace much of human wage labor over the next 20 years and push many of us, and myself almost certainly included, into premature retirement. I'm personally concerned that in a few years, I'll find my software proficiency to be as useful as my chess proficiency today is useful to Stockfish. I am afraid of a massive social upheaval both for myself and my family, and for society…

Here “much of” is doing the heavy lifting. Are you willing to commit to a percentage or a range? I work at an insurance company and I can’t see AI replacing even 10% of the employees here. Too much of what we do is locked up in decades-old proprietary databases that cannot be replaced for legal reasons. We still rely on paper mail for a huge amount of communication with policyholders. The decisions we make on a daily…

Yes, absolutely willing to commit. I can't find a single reliable source, but from what I gather, over 70% of people in the West do "pure knowledge work", which doesn't include any embodied actuvities. I am happy to put my money that these jobs will start being fully taken over by AI rapidly soon (if they aren't already), and that by 2035, less than 50% of us will have a job that doesn't require "being there".

And regarding your example of an insurance company, I'm not sure about that industry, but seeing the transformation of banking over the last decade to fully digital providers like Revolut, I would expect similar disruption there.

Post reply on HN