I like to imagine that the number of consumed tokens before a solution is found is a proxy for how difficult a problem is, and it looks like Opus 4.6 consumed around 250k tokens. That means that a tricky React refactor I did earlier today at work was about half as hard as an open problem in mathematics! :)
You might be joking, but you're probably also not that far off from reality. I think more people should question all this nonsense about AI "solving" math problems. The details about human involvement are always hazy and the significance of the problems are opaque to most. We are very far away from the sensationalized and strongly implied idea that we are doing something miraculous here.
Epoch confirms GPT5.4 Pro solved a frontier math open problem
21–30 of 744 posts
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#22As someone with only passing exposure to serious math, this section was by far the most interesting to me: > The author assessed the problem as follows. > [number of mathematicians familiar, number trying, how long an expert would take, how notable, etc] How reliably can we know these things a-priori? Are these mostly guesses? I don't mean to diminish the value of guesses; I'm curious how reliable these kinds of gues…
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#23Can an AI pose an frontier math problem that is of any interest to mathematicians?
I would guess 1) AI can solve frontier math problems and 2) can pose interesting/relevant math problems together would be an "oh shit" moment. Because that would be true PhD level research.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#24No denial at this point, AI could produce something novel, and they will be doing more of this moving forward.
Not sure if AI can have clever or new ideas, it still seems to be it combines existing knowledge and executes algoritms. I am not necessarily saying humans do something different either, but I have yet to see a novel solution from an AI that is not simply an extrapolation of current knowledge.
Sometimes just having the time/compute to explore the available space with known knowledge is enough to produce something unique.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#25Earlier quoted context omitted.
You might be joking, but you're probably also not that far off from reality. I think more people should question all this nonsense about AI "solving" math problems. The details about human involvement are always hazy and the significance of the problems are opaque to most. We are very far away from the sensationalized and strongly implied idea that we are doing something miraculous here.
I am kind of joking, but I actually don't know where the flaw in my logic is. It's like one of those math proofs that 1 + 1 = 3. If I were to hazard a guess, I think that tokens spent thinking through hard math problems probably correspond to harder human thought than tokens spend thinking through React issues. I mean, LLMs have to expend hundreds of tokens to count the number of r's in strawberry. You can't tell me…
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#26Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#27Earlier quoted context omitted.
You might be joking, but you're probably also not that far off from reality. I think more people should question all this nonsense about AI "solving" math problems. The details about human involvement are always hazy and the significance of the problems are opaque to most. We are very far away from the sensationalized and strongly implied idea that we are doing something miraculous here.
I am kind of joking, but I actually don't know where the flaw in my logic is. It's like one of those math proofs that 1 + 1 = 3. If I were to hazard a guess, I think that tokens spent thinking through hard math problems probably correspond to harder human thought than tokens spend thinking through React issues. I mean, LLMs have to expend hundreds of tokens to count the number of r's in strawberry. You can't tell me…
1. Knowing how to state the problem. Ie, go from the vague problem of "I don't like this, but I do like this", to the more specific problem of "I desire property A". In math a lot of open problems are already precisely stated, but then the user has to do the work of _understanding_ what the precise stating is.
2. Verifying that the proposed solution actually is a full solution.
This math problem actually illustrates them both really well to me. I read the post, but I still couldn't do _either_ of the steps above, because there's a ton of background work to be done. Even if I was very familiar with the problem space, verifying the solution requires work -- manually looking at it, writing it up in coq, something like that. I think this is similar to the saying "it takes 10 years to become an overnight success"
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#28I like to imagine that the number of consumed tokens before a solution is found is a proxy for how difficult a problem is, and it looks like Opus 4.6 consumed around 250k tokens. That means that a tricky React refactor I did earlier today at work was about half as hard as an open problem in mathematics! :)
You might be joking, but you're probably also not that far off from reality. I think more people should question all this nonsense about AI "solving" math problems. The details about human involvement are always hazy and the significance of the problems are opaque to most. We are very far away from the sensationalized and strongly implied idea that we are doing something miraculous here.
Not really. You're just in denial and are not really all that interested in the details. This very post has the transcript of the chat of the solution.
Re: Epoch confirms GPT5.4 Pro solved a frontier math open problem
#29I feel like there’s a fork in our future approaching where we’ll either blossom into a paradise for all or live under the thumb of like 5 immortal VCs
Change is always hard, even if it will be good in 20 years, the transitions are always tough.
Hoping that won't be the case with AI but we may need some major societal transformations to prevent it.