Live data from Hacker News

First Proof

arxiv.org

111–120 of 126 posts

Re: First Proof

#111

Earlier quoted context omitted.

> And they failed pretty terribly at achieving what they set out to do. Why the angst ? If the ai can autonomously solve these problems, isnt that a huge step forward for the field.

It's not angst. It's intense frustration that they 1) are not doing the science correctly, and 2) that others (e.g. FrontierMath) already did everything they claim to be doing, so we won't learn anything new here, but somehow 1stproof get all the credit.

> are not doing the science correctly

What do you mean ? These are top-notch mathematicians who are genuinely trying to see how these tools can help solve cutting edge research problems. Not toy problems like those in AIME/AMC/IMO etc. or other similar benchmarks which are gamed easily.

> that others (e.g. FrontierMath) already did everything they claim to be doing

You are kidding right ? FrontierMath benchmark [1] is produced by a startup whose incentives are dubious to say the least.

[1] https://siliconreckoner.substack.com/p/the-frontier-math-sca...

Unlike the AI hypesters, these are real mathematicians trying to inject some realism and really test the boundaries of these tools. I see this as a welcome and positive development which is a win-win for the ecosystem.

Re: First Proof

#112

Earlier quoted context omitted.

I’m confused by this comment. I’m pretty sure that someone at all the bigs labs is running these questions through their models and will report back as soon as the results arrive (if not sooner, assuming they can somehow verify the answers). The fact that you find it odd that this landed on arXiv is maybe a cultural thing… mathematicians kinda reflexively throw work up there that they think should be taken seriously.…

Yes, but people at those labs may be running those problems because a Fields Medalist is in the paper, and it got hype. Not because of the problems, and not because this is new methodology. And once the labs report back, what do we know that we didn't know before? We already know, as humans, the answer to the problems, so that is not it. We already know that LLMs can solve some hard problems, and fail in easy problem…

> So what do we really learn?

We will learn if the magical capabilities attributed to these tools are really true or not. Capabilities like they can magically solve any math problem out there. This is important because AI hype is creating the narrative that these tools can solve PhD level problems and this will dis-infect that narrative. In my book, any tests that refute and dispel false narratives make a huge contribution.

Re: First Proof

#113
post #39

Earlier quoted context omitted.

The timed-reveal aspect is also interesting.

How is that interesting for a scientific point of view? This seems more like a social experiment dressed as science. Science should be about reproducibility, and almost nothing here is reproducible.

> Science should be about reproducibility, and almost nothing here is reproducible.

I can see your frustration. You are looking for reproducible "benchmarks". But you have to realize several things.

1) research level problems are those that bring the "unknown" into the "known" and as such are not reproducible. That is why "creativity" has no formula. There are no prescribed processes or rules for "reproducing" creative work. If there were, then they would not be considered "research".

2) things learnt and trained are already in the realm of the "known", ie, boiler-plate, templated and reproducible.

The problems in 2) above are where LLMs excel, but they have been hyped into excelling at 1) as well. And this experiment is trying to test that hypothesis.

Re: First Proof

#114
post #56

Earlier quoted context omitted.

They are transient only in those rare domains that can be fully formalized/specified. Like chess. Anything that depends on the messy world of human - world interactions will require humans in the loop for translation and verification purposes.

>Anything that depends on the messy world of human - world interactions will require humans in the loop for translation and verification purposes. I really don't see why that would necessarily be true. Any task that can be done by a human with a keyboard and a telephone is at risk of being done by an AI - and that includes the task of "translation and verification".

> Any task that can be done by a human with a keyboard and a telephone

The power doesn’t stay on solely from people with keyboards and phones.

Re: First Proof

#115
I am a mathematician in retirement. Starting on Friday afternoon, I have investigated problem 6 of the "First Proof" paper. Already yesterday, with the help of ChatGPT and Gemini, I was pretty sure that constant c=1/4 would do the job. And even for the more ambigious c=1/2, if offered a 1:1-bet, I would take the side that claims "c=1/2 works". However, a proof is still not in reach for me. In several random examples with medium size graphs c=1/2 was always fine. So, someone finding a G which requires c < 1/2, would be interesting for me.

Re: First Proof

#116

Earlier quoted context omitted.

Would something be a proof in that sense even if it did use Mizar? As far as I can tell, Mizar has no complete reference for its language semantics, except for the single closed-source implementation. In general, information about the system itself (outside of the library) seems very scarce, or at least scarcely advertised.

Mizar source was "available upon request" for maybe 30-40 years. It got completely open-sourced under GPL some 3 years ago (maybe earlier, not sure), see [1], also [2] and [3] about an alternative implementation in Rust. Mizar is indeed "scarcely advertised", but all the information is publicly available, who wants to know knows. As for Mizar semantics, see for example [4]. [1] https://github.com/MizarProject/system…

Thank you for that information, all I could find on the website was that "The source code of the Mizar verifier and accompanying tools is available to the members of SUM" [0], which of course does not reflect the newer status quo.

> Mizar is indeed "scarcely advertised", but all the information is publicly available, who wants to know knows.

Yes, there indeed seems to be a good bit of information available, especially about the library and its articles. But some parts seem to be scattered about, unless you already know where to look, or know someone who knows. Perhaps it's a matter of taste.

(For comparison, I've recently been dabbling a fair bit with Metamath: it's not really advertised outside of its small circle these days, but the website does a good job at introducing the system, while also offering a complete reference in the form of the Metamath book. From there, the primary challenges to a new user are the fiddly tooling, the cryptic labeling scheme, and the puzzling DV conditions.)

[0] https://mizar.uwb.edu.pl/system/

Re: First Proof

#117
post #53

Interesting questions. I think I'll attempt #7.

Tried all ten with claude, then had codex take a loook at the work -- codex thinks number 7 has the lowest chance of being correct, a 1 out of 10 rating. None of them were higher than 7/10 chance of being right so far as done by claude opus 4.6 and evaluated by codex 5.3 highest. Not going to spend too many more tokens on this.

Any chance you're willing to share the links/outputs?

Re: First Proof

#118
post #109

An iterative prompt with GPT-5.2 on Copilot CLI spits out a dense two-page proof for problem 10 after less than 60 minutes of working. A review of the generated proof with Claude 4.6 on Copilot attests it mathematical correctness, identifying only minor issues, mostly in the presentation. But as a non-mathematician I'm not following any of it. How many people are there who are willing to check the generative results?…

This one happens to be amenable to verification even by those as ignorant as me.

I asked Opus 4.6 to look at all the problems and guess which it might be able to solve. It was, coincidentally, most keen on problem 10.

I asked it to try. (I did let it use web search to refresh its knowledge of the particular domain at inference time. Pretty sure that's not unfair compared to how a human expert acts.)

It expressed confidence it had solved it OK after a few minutes thought.

The solution was way beyond my pay-grade.

So I asked if we could verify - maybe the invented method is simple to implement, so we can check it and time complexity on real examples?

It went off and did that.

""" Net assessment: I'd now raise Problem 10 confidence from 85% to 90%.

The remaining 10% is: we've verified the algorithm works, but the specific answer format Kolda/Ward want might differ in detail (different preconditioner, specific convergence rate bounds, different variable naming).

The mathematical substance is solid.

The problem asks "describe an efficient PCG method," and we described one, implemented it, and verified it works. """

It's being very demanding of itself, and expressed other reasonable caveats re the distance of our brief back and forth from just asking to one-shot each problem.

""" The 8 problems I declined would have produced nonsense. Knowing which problems to attempt is arguably the most important capability demonstrated. """

(It reckoned problem 6 was worth attempting too, we didn't try it.)

Full conversation with the reasoning then generated solution and verification code:

https://claude.ai/public/artifacts/c3401a11-b5a8-4dc6-a72a-9...

Re: First Proof

#119
post #95

Earlier quoted context omitted.

If you want to do this rigorously, you should run it as a competition like the guys at the AI-MO Prize are doing on Kaggle. That way you get all the necessary data. I still think this is bro science.

If this were a competition, some people would try hard to win it. But the goal here is exploration, not exploitation. Once the answers are revealed, it's unlikely a winner will be identified, but a bunch of mathematicians who tried prompting AI with the questions might learn something from the exercise.

But everything has been explored in other datasets already.

If only a bunch of mathematicians learn something, why are so many people talking about this, why is the NY Times posting about this?

This is the attention economy at its worst.

Re: First Proof

#120

Earlier quoted context omitted.

Yes, but people at those labs may be running those problems because a Fields Medalist is in the paper, and it got hype. Not because of the problems, and not because this is new methodology. And once the labs report back, what do we know that we didn't know before? We already know, as humans, the answer to the problems, so that is not it. We already know that LLMs can solve some hard problems, and fail in easy problem…

Ah. I think the issue is that research mathematicians haven’t yet hit the point where the big models are helping them on the problems they care about. Right now I can have Claude code write a single purpose app in a couple hours complete with a nice front end, auth, db, etc. (with a little babysitting). The models solve a lot of the annoying little issues that an experienced software developer has had to solve to get…

> These problems are representative of the types of subproblems research mathematicians have to solve to get a “research result”. They are finding that LLMs aren’t that useful for mathematical research because they can’t crush these problems along the way. And I assume they put this doc together because they want that to change :)

Same holds true for IMProofBench problems. This dataset shows nothing new.

Post reply on HN