Live data from Hacker News

Learning more about Claude's mathematical capabilities

anthropic.com

141–150 of 185 posts

Re: Learning more about Claude's mathematical capabilities

#141
post #35

prompt engineering 2025: you are an expert programmer, use industry best practices, test driven development and use modularity and abstraction to anticipate future features, … prompt engineering 2026: i believe in you

Yes, one of the things I find hardest about using Claude code for Mathematics is keeping negativity out of the notes and memory. First of all, it will convince itself that a task is just too hard and find excuses not to try hard enough. Then when it struggles on something it loves to write down confusing notes about things which it believes cases the problem. Then next iteration it reads its own note, misinterprets i…

Yet a significant percentage of people simultaneously believe it "reasoned" its way to novel mathematical proofs. Something doesn't jive

Re: Learning more about Claude's mathematical capabilities

#142
post #69

I wonder why we have yet to see more systematic exploration of Math. Anthropic describes that Claude identified a set of possibilities and then explored them using sub-agents. The human saying "I believe in you" could literally just be something along lines of a harness with a /goal loop. We all identify this as absurd because... it's so lacking in rigor despite making major progress. What if we just applied a little…

I wouldn't be surprised if half these proofs turn out to be well crafted hallucinations, barring of course the ones actually verified in Lean

Re: Learning more about Claude's mathematical capabilities

#145

Earlier quoted context omitted.

$2M TC. Job: AI cheerleader.

Have you seen how much college football coaches make? The world collectively spends trillions training, and entertaining and then watching people kick and hit balls around but somehow working with AI is preposterous for less money.

Asking the AI to believe in itself is the absurd part.

I don’t understand why people think describing something in some ostensibly dismissive way constitutes making a point. But some people spend their time banging keys with letters on them (or not with letters on them) into input fields and then pressing Return, so I guess some subset of those people will do that.

Re: Learning more about Claude's mathematical capabilities

#146
post #43

> Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”). This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress. I remain delighted at how absurd our current timeline has become.

"Any sufficiently advanced technology is indistinguishable from magic." -- Arthur C. Clarke's Third Law

"Sometimes, magic is just someone spending more time on something than anyone else might reasonably expect." -- Teller (of Penn & Teller)

"Sometimes, any sufficiently advanced technology is just spending more time on something than anyone else might reasonably expect." -- an LLM's original thought, probably

Re: Learning more about Claude's mathematical capabilities

#147
It's very very important to note that there was an existing 2025 arxiv preprint that had a >66% proof assuming some weak condition, and this result removes that weak condition.

It didn't do this whole 41.6->67.2% jump by itself, humans had done most of the work and it came in at the end and found a way to remove the condition. Impressive, but not as massively impressive as when it sounds like it did the jump by itself.

This isn't goalpost moving, it's clarifying what exactly happened bc at first I thought it had made the jump by itself. The blog post is written in a technically correct, but misleading way where it takes credit for the whole jump.

Re: Learning more about Claude's mathematical capabilities

#148
post #132

Earlier quoted context omitted.

Yes, one of the things I find hardest about using Claude code for Mathematics is keeping negativity out of the notes and memory. First of all, it will convince itself that a task is just too hard and find excuses not to try hard enough. Then when it struggles on something it loves to write down confusing notes about things which it believes cases the problem. Then next iteration it reads its own note, misinterprets i…

interesting! it sounds like this might benefit from a ui that helps to edit/re-write the history I also think this would make sense for programming but there it is a bit harder to justify the effort when you are working on something that really matters this can make the difference though ty for sharing!

(not the above poster but) I set up a local workspace for it to read/write from, so that edit is effectively just VSCode/vim. Unfortunately, it seems more verbose when writing out to file.

Re: Learning more about Claude's mathematical capabilities

#149
post #57
post #9

Since they say that this is from an unreleased research version of Claude: I wonder if at some point Anthropic and OpenAI will start delaying the release of their models intentionally so they can reap the benefits from the models in, for example, mathematics, medicine, physics, and other fields. Just as an example, imagine if your model were capable of proving P = NP, or if your model could cure diseases. Would you r…

>From these companies' standpoint, I think they would choose the latter. Ever since these things came about I've wondered why they haven't been doing this the whole time. If they've got the "do-anything" robot and can scale a billion of them, why aren't they creating a Do-Everything conglomerate that disrupts every possible industry with zero/negligible labor costs? The only answer I've come up with is that they stil…

Because, like 98% of people in this space, you don't mention or even consider cost. Improving the lower bound of Riemann is impressive, but how impressive would it remain if it was announced that training and inference cost $1 billion dollars?

Not as much, I predict

Re: Learning more about Claude's mathematical capabilities

#150
post #120

I'm not sure what's crazier: AI improving a lower bound on RH, or AI improving a lower bound on RH and it not even making the front page of HN.

Right? I just learned about this and searched it on hacker news wondering why I didn't see it earlier. Any human mathematician would be thrilled to prove a result like this, and it's not even big news anymore that an LLM can do it.

Where did you find the news?

So far I only found decent discussions about it in Chinese.

https://www.zhihu.com/question/2070336637360518307/answer/20...

Post reply on HN