Live data from Hacker News

First Proof

arxiv.org

81–90 of 126 posts

Re: First Proof

#81

I'm a mathematician relying heavily on AI as an association engine of massive scope, to organize and expand my thoughts. One doesn't get best results by "testing" AI. A surfboard is also an amazing tool, but there's more to operating one than telling it which way to go. Many people want self-driving cars so they can drink in the back seat watching movies. They'll find their jobs replaced by AI, with a poor quality of…

> Anthropic is successfully coding Claude using Claude.

Claude is one of the buggiest pieces of shit I have ever used. They had to BUY the creators of bun to fix the damn thing. It is not a good example of your thesis.

Re: First Proof

#82

I'm a mathematician relying heavily on AI as an association engine of massive scope, to organize and expand my thoughts. One doesn't get best results by "testing" AI. A surfboard is also an amazing tool, but there's more to operating one than telling it which way to go. Many people want self-driving cars so they can drink in the back seat watching movies. They'll find their jobs replaced by AI, with a poor quality of…

> At the same time, Anthropic is successfully coding Claude using Claude. Is that why everyone keeps complaining about the quality getting worse?

I think that’s more about model performance degrading due to less computational resources being assigned to them over time.

Re: First Proof

#83

I'm a mathematician relying heavily on AI as an association engine of massive scope, to organize and expand my thoughts. One doesn't get best results by "testing" AI. A surfboard is also an amazing tool, but there's more to operating one than telling it which way to go. Many people want self-driving cars so they can drink in the back seat watching movies. They'll find their jobs replaced by AI, with a poor quality of…

> Anthropic is successfully coding Claude using Claude. Claude is one of the buggiest pieces of shit I have ever used. They had to BUY the creators of bun to fix the damn thing. It is not a good example of your thesis.

You and the GP are conflating Claude, the company or its flagship model Claude Opus, with Claude Code, a state of the art coding assistant that has admittedly a slow and buggy React-based TUI (output quality is still very competitive)

Re: First Proof

#85
post #67

Earlier quoted context omitted.

Centaurs are a transient phenomenon. In chess, the era of centaur supremacy lasted only about a decade before computers alone eclipsed human+computer. The same will be true in every other discipline. You can surf the wave, but sooner or later, the wave will come crashing down.

Last I heard, which was last year, human + computer still beat either by themselves. You got a link about what's changed?

You're the one claiming "Last I heard" so you're the one who owes a link.

Re: First Proof

#86

These are very serious research level math questions. They are not “Erdős style” questions; they look more like problems or lemmas that I encountered while doing my PhD. Things that don’t make it into the papers but were part of an interesting diversion along the way. It seems likely that PhD students in the subfields of the authors are capable of solving these problems. What makes them interesting is that they seem…

So these are like those problems that are “left for the reader”?

No, results in a paper are identified to be "left for the reader" because they are thought to be straightforward to the paper's audience. These are chosen because they are novel. I didn't see any reason to think they are easier than the main results, just maybe not of as much interest.

Re: First Proof

#87

Earlier quoted context omitted.

Very serious for mathematicians - not for ML researchers. If the paper would not have had the AI spin, would those 10 questions still have been interesting? It seems to me that we have here a paper that is solely interesting because of the AI spin -- while at the same time this AI spin is really poorly executed from the point of AI research, where this should be a blog post at most, not an arXiv preprint.

I’m confused by this comment. I’m pretty sure that someone at all the bigs labs is running these questions through their models and will report back as soon as the results arrive (if not sooner, assuming they can somehow verify the answers). The fact that you find it odd that this landed on arXiv is maybe a cultural thing… mathematicians kinda reflexively throw work up there that they think should be taken seriously.…

Yes, but people at those labs may be running those problems because a Fields Medalist is in the paper, and it got hype.

Not because of the problems, and not because this is new methodology.

And once the labs report back, what do we know that we didn't know before? We already know, as humans, the answer to the problems, so that is not it. We already know that LLMs can solve some hard problems, and fail in easy problems, so that is not it either.

So what do we really learn?

Re: First Proof

#88

Earlier quoted context omitted.

How is that interesting for a scientific point of view? This seems more like a social experiment dressed as science. Science should be about reproducibility, and almost nothing here is reproducible.

Deepmind’s Nobel Prize was primarily for its performance in CASP which is pretty much exactly this. Labs solve structures of proteins, but don’t publish them until after all the computational teams predict structures. So I’m not sure where you’re coming from claiming that this isn’t scientific.

It wasn't like this in any way.

CASP relies on a robust benchmark (not just 10 random proteins), and has clear participation criteria, objective metrics how the eval plays out, etc.

So I stand by my claim: This isn't scientific. If CASP is Japan, a highly organized & civilized society, this is a banana republic.

Re: First Proof

#89

Earlier quoted context omitted.

How is that interesting for a scientific point of view? This seems more like a social experiment dressed as science. Science should be about reproducibility, and almost nothing here is reproducible.

Reproducibility is just one aspect of science, logic + reasoning from principles and data is the major aspect. There are some experiments which cannot be carried out more than once.

> There are some experiments which cannot be carried out more than once

Yes, in which case a very detailed methodology is required: which hardware, runtimes, token counts etc.

This does none of that.

Re: First Proof

#90

Earlier quoted context omitted.

Looks like very sloppy research.

I don't think it's that serious...it's an interesting experiment that assumes people will take it in good faith. The idea is also of course to attach the transcript log and how you prompted the LLM so that anyone can attempt to reproduce if they wish.

If you want to do this rigorously, you should run it as a competition like the guys at the AI-MO Prize are doing on Kaggle.

That way you get all the necessary data.

I still think this is bro science.

Post reply on HN