Live data from Hacker News

First Proof

arxiv.org

61–70 of 126 posts

Re: First Proof

#61
post #51

Earlier quoted context omitted.

Hey, sorry, totally out of context but I've always wanted to ask about the username. I keep reading it as "yoruba" in my mind. What does it mean, if I'm not being indiscreet?

You're not the first to have wondered: https://news.ycombinator.com/item?id=20730027

Well, now that I read that comment I remembered having read it before. My mind is going.

Re: First Proof

#62
post #20

Earlier quoted context omitted.

The authors mention that before publications they tested these questions on Gemini and GPT, so they have been available to the two biggest players already; they have a head start.

Looks like very sloppy research.

I don't think it's that serious...it's an interesting experiment that assumes people will take it in good faith. The idea is also of course to attach the transcript log and how you prompted the LLM so that anyone can attempt to reproduce if they wish.

Re: First Proof

#63
post #53

Interesting questions. I think I'll attempt #7.

Tried all ten with claude, then had codex take a loook at the work -- codex thinks number 7 has the lowest chance of being correct, a 1 out of 10 rating. None of them were higher than 7/10 chance of being right so far as done by claude opus 4.6 and evaluated by codex 5.3 highest. Not going to spend too many more tokens on this.

I don't think either of these are the best choices for this. Chatgpt 5.2 pro and gemini 3 pro deep thinking I believe are the strongest LLMs at "pure thought", i.e. things like mathematical reasoning.

Re: First Proof

#64

Can someone explain how this would work? > the answers are known to the authors of the questions but will remain encrypted for a short time. Ok. But humans may be able to solve the problems too. What prevents Anthropic or OpenAI from hiring mathematicians, have them write the proof and pass it off as LLM written? I'm not saying that's what they'll do. But shouldn't the paper say something about how they're going to v…

Because LLMs are deterministic, they could provide the model files, prompt, and seed used.

Re: First Proof

#65

I'm a mathematician relying heavily on AI as an association engine of massive scope, to organize and expand my thoughts. One doesn't get best results by "testing" AI. A surfboard is also an amazing tool, but there's more to operating one than telling it which way to go. Many people want self-driving cars so they can drink in the back seat watching movies. They'll find their jobs replaced by AI, with a poor quality of…

What a beautifully articulated take!

Re: First Proof

#66
post #34

Earlier quoted context omitted.

Paper not about benchmarking or ML research is bad from the perspective of benchmarking. Not exactly a shocker. The authors themselves literally state: "Unlike other proposed math research benchmarks (see Section 3), our question list should not be considered a benchmark in its current form"

On the website https://1stproof.org/#about they claim: "This project represents our preliminary efforts to develop an objective and realistic methodology for assessing the capabilities of AI systems to autonomously solve research-level math questions." Sounds to me to be a benchmark in all but a name. And they failed pretty terribly at achieving what they set out to do.

> And they failed pretty terribly at achieving what they set out to do.

Why the angst ? If the ai can autonomously solve these problems, isnt that a huge step forward for the field.

Re: First Proof

#67

I'm a mathematician relying heavily on AI as an association engine of massive scope, to organize and expand my thoughts. One doesn't get best results by "testing" AI. A surfboard is also an amazing tool, but there's more to operating one than telling it which way to go. Many people want self-driving cars so they can drink in the back seat watching movies. They'll find their jobs replaced by AI, with a poor quality of…

Centaurs are a transient phenomenon. In chess, the era of centaur supremacy lasted only about a decade before computers alone eclipsed human+computer. The same will be true in every other discipline. You can surf the wave, but sooner or later, the wave will come crashing down.

Last I heard, which was last year, human + computer still beat either by themselves. You got a link about what's changed?

Re: First Proof

#68

I'm a mathematician relying heavily on AI as an association engine of massive scope, to organize and expand my thoughts. One doesn't get best results by "testing" AI. A surfboard is also an amazing tool, but there's more to operating one than telling it which way to go. Many people want self-driving cars so they can drink in the back seat watching movies. They'll find their jobs replaced by AI, with a poor quality of…

> At the same time, Anthropic is successfully coding Claude using Claude.

Is that why everyone keeps complaining about the quality getting worse?

Re: First Proof

#69
post #56

Earlier quoted context omitted.

Centaurs are a transient phenomenon. In chess, the era of centaur supremacy lasted only about a decade before computers alone eclipsed human+computer. The same will be true in every other discipline. You can surf the wave, but sooner or later, the wave will come crashing down.

They are transient only in those rare domains that can be fully formalized/specified. Like chess. Anything that depends on the messy world of human - world interactions will require humans in the loop for translation and verification purposes.

How is chess not fully specified?

Re: First Proof

#70
post #39

Earlier quoted context omitted.

The timed-reveal aspect is also interesting.

How is that interesting for a scientific point of view? This seems more like a social experiment dressed as science. Science should be about reproducibility, and almost nothing here is reproducible.

Reproducibility is just one aspect of science, logic + reasoning from principles and data is the major aspect.

There are some experiments which cannot be carried out more than once.

Post reply on HN