Earlier quoted context omitted.
Hey, sorry, totally out of context but I've always wanted to ask about the username. I keep reading it as "yoruba" in my mind. What does it mean, if I'm not being indiscreet?
You're not the first to have wondered: https://news.ycombinator.com/item?id=20730027
First Proof
61–70 of 126 posts
Re: First Proof
#62Earlier quoted context omitted.
The authors mention that before publications they tested these questions on Gemini and GPT, so they have been available to the two biggest players already; they have a head start.
Looks like very sloppy research.
Re: First Proof
#63Interesting questions. I think I'll attempt #7.
Tried all ten with claude, then had codex take a loook at the work -- codex thinks number 7 has the lowest chance of being correct, a 1 out of 10 rating. None of them were higher than 7/10 chance of being right so far as done by claude opus 4.6 and evaluated by codex 5.3 highest. Not going to spend too many more tokens on this.
Re: First Proof
#64Can someone explain how this would work? > the answers are known to the authors of the questions but will remain encrypted for a short time. Ok. But humans may be able to solve the problems too. What prevents Anthropic or OpenAI from hiring mathematicians, have them write the proof and pass it off as LLM written? I'm not saying that's what they'll do. But shouldn't the paper say something about how they're going to v…
Re: First Proof
#65I'm a mathematician relying heavily on AI as an association engine of massive scope, to organize and expand my thoughts. One doesn't get best results by "testing" AI. A surfboard is also an amazing tool, but there's more to operating one than telling it which way to go. Many people want self-driving cars so they can drink in the back seat watching movies. They'll find their jobs replaced by AI, with a poor quality of…
Re: First Proof
#66Earlier quoted context omitted.
Paper not about benchmarking or ML research is bad from the perspective of benchmarking. Not exactly a shocker. The authors themselves literally state: "Unlike other proposed math research benchmarks (see Section 3), our question list should not be considered a benchmark in its current form"
On the website https://1stproof.org/#about they claim: "This project represents our preliminary efforts to develop an objective and realistic methodology for assessing the capabilities of AI systems to autonomously solve research-level math questions." Sounds to me to be a benchmark in all but a name. And they failed pretty terribly at achieving what they set out to do.
Why the angst ? If the ai can autonomously solve these problems, isnt that a huge step forward for the field.
Re: First Proof
#67I'm a mathematician relying heavily on AI as an association engine of massive scope, to organize and expand my thoughts. One doesn't get best results by "testing" AI. A surfboard is also an amazing tool, but there's more to operating one than telling it which way to go. Many people want self-driving cars so they can drink in the back seat watching movies. They'll find their jobs replaced by AI, with a poor quality of…
Centaurs are a transient phenomenon. In chess, the era of centaur supremacy lasted only about a decade before computers alone eclipsed human+computer. The same will be true in every other discipline. You can surf the wave, but sooner or later, the wave will come crashing down.
Re: First Proof
#68I'm a mathematician relying heavily on AI as an association engine of massive scope, to organize and expand my thoughts. One doesn't get best results by "testing" AI. A surfboard is also an amazing tool, but there's more to operating one than telling it which way to go. Many people want self-driving cars so they can drink in the back seat watching movies. They'll find their jobs replaced by AI, with a poor quality of…
Is that why everyone keeps complaining about the quality getting worse?
Re: First Proof
#69Earlier quoted context omitted.
Centaurs are a transient phenomenon. In chess, the era of centaur supremacy lasted only about a decade before computers alone eclipsed human+computer. The same will be true in every other discipline. You can surf the wave, but sooner or later, the wave will come crashing down.
They are transient only in those rare domains that can be fully formalized/specified. Like chess. Anything that depends on the messy world of human - world interactions will require humans in the loop for translation and verification purposes.
Re: First Proof
#70Earlier quoted context omitted.
The timed-reveal aspect is also interesting.
How is that interesting for a scientific point of view? This seems more like a social experiment dressed as science. Science should be about reproducibility, and almost nothing here is reproducible.
There are some experiments which cannot be carried out more than once.