Live data from Hacker News

First Proof

arxiv.org

51–60 of 126 posts

Re: First Proof

#51
post #11

Earlier quoted context omitted.

This is not a benchmark. They just want to give people the opportunity to try their hand at solving novel questions with AI and see what happens. If an AI company pulls a solution out of their hat that cannot be replicated with the products they make available to ordinary people, that's hardly worth bragging about and in any case it's not the point of the exercise.

Hey, sorry, totally out of context but I've always wanted to ask about the username. I keep reading it as "yoruba" in my mind. What does it mean, if I'm not being indiscreet?

You're not the first to have wondered: https://news.ycombinator.com/item?id=20730027

Re: First Proof

#52
post #9

Earlier quoted context omitted.

The abstract of the article is very short, and seems pretty clear to both of your questions. This is what is special about them: > a set of ten math questions which have arisen naturally in the research process of the authors. The questions had not been shared publicly until now; I.e. these are problems of some practical interest, not just performative/competitive maths. And this is what is know about the solutions:…

> these are problems of some practical interest, not just performative/competitive maths. FrontierMath did this a year ago. Where is the novelty here? > a solution is known, but is guaranteed to not be in the training set for any AI. Wrong, as the questions were poses to commercial AI models and they can solve them. This paper violates basic benchmarking principles.

> Wrong, as the questions were poses to commercial AI models and they can solve them.

Why does this matter? As far as I can tell, because the solution is not known this only affects the time constant (i.e. the problems were known for longer than a week). It doesn't seem that I should care about that.

Re: First Proof

#53

Interesting questions. I think I'll attempt #7.

Tried all ten with claude, then had codex take a loook at the work -- codex thinks number 7 has the lowest chance of being correct, a 1 out of 10 rating. None of them were higher than 7/10 chance of being right so far as done by claude opus 4.6 and evaluated by codex 5.3 highest.

Not going to spend too many more tokens on this.

Re: First Proof

#54

Earlier quoted context omitted.

Centaurs are a transient phenomenon. In chess, the era of centaur supremacy lasted only about a decade before computers alone eclipsed human+computer. The same will be true in every other discipline. You can surf the wave, but sooner or later, the wave will come crashing down.

Agreed. But I don't think the time scale will be similar. Chess is relatively simple in comparison, as complex as it is.

On the other hand, chess is not very financially rewarding. IBM put some money into it for marketing briefly, but that’s probably equal to about five minutes of spend from the current crop of LLM companies.

Re: First Proof

#56

I'm a mathematician relying heavily on AI as an association engine of massive scope, to organize and expand my thoughts. One doesn't get best results by "testing" AI. A surfboard is also an amazing tool, but there's more to operating one than telling it which way to go. Many people want self-driving cars so they can drink in the back seat watching movies. They'll find their jobs replaced by AI, with a poor quality of…

Centaurs are a transient phenomenon. In chess, the era of centaur supremacy lasted only about a decade before computers alone eclipsed human+computer. The same will be true in every other discipline. You can surf the wave, but sooner or later, the wave will come crashing down.

They are transient only in those rare domains that can be fully formalized/specified. Like chess. Anything that depends on the messy world of human - world interactions will require humans in the loop for translation and verification purposes.

Re: First Proof

#57

No, this is not a proof because not using Mizar ;-) https://mizar.uwb.edu.pl/

Would something be a proof in that sense even if it did use Mizar? As far as I can tell, Mizar has no complete reference for its language semantics, except for the single closed-source implementation. In general, information about the system itself (outside of the library) seems very scarce, or at least scarcely advertised.

Re: First Proof

#58
post #39

Earlier quoted context omitted.

Very serious for mathematicians - not for ML researchers. If the paper would not have had the AI spin, would those 10 questions still have been interesting? It seems to me that we have here a paper that is solely interesting because of the AI spin -- while at the same time this AI spin is really poorly executed from the point of AI research, where this should be a blog post at most, not an arXiv preprint.

The timed-reveal aspect is also interesting.

How is that interesting for a scientific point of view? This seems more like a social experiment dressed as science.

Science should be about reproducibility, and almost nothing here is reproducible.

Re: First Proof

#59

Earlier quoted context omitted.

> these are problems of some practical interest, not just performative/competitive maths. FrontierMath did this a year ago. Where is the novelty here? > a solution is known, but is guaranteed to not be in the training set for any AI. Wrong, as the questions were poses to commercial AI models and they can solve them. This paper violates basic benchmarking principles.

> Wrong, as the questions were poses to commercial AI models and they can solve them. Why does this matter? As far as I can tell, because the solution is not known this only affects the time constant (i.e. the problems were known for longer than a week). It doesn't seem that I should care about that.

Because the companies have the data and can solve them -- so providing the question to a company with the necessary manpower, one cannot guarantee anymore that the solution is not known, and not contained in the training sample.

Re: First Proof

#60
post #56

Earlier quoted context omitted.

Centaurs are a transient phenomenon. In chess, the era of centaur supremacy lasted only about a decade before computers alone eclipsed human+computer. The same will be true in every other discipline. You can surf the wave, but sooner or later, the wave will come crashing down.

They are transient only in those rare domains that can be fully formalized/specified. Like chess. Anything that depends on the messy world of human - world interactions will require humans in the loop for translation and verification purposes.

From a human, to a centaur, to a pegasus, as it were.
Post reply on HN