Live data from Hacker News

DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]

github.com

51–53 of 53 posts

Re: DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]

#51

Earlier quoted context omitted.

For one thing, it's not a real score; they judged the results themselves and Putnam judges are notoriously tough. There was not a single 8 on the problem they claim partial credit for (or any partial credit above a 2) amongst the top 500 humans. https://kskedlaya.org/putnam-archive/putnam2024stats.html . For another thing, the 2024 Putnam problems are in their RL data. Also, it's very unclear how these competitions c…

What do other models trained on the same problems score? What about if they are RL'd to not reproduce things word for word? Why do you think that the 2024 Putnam programs that they used to test were in the training data? /? "Art of Problem Solving" Putnam https://www.google.com/search?q=%22Art+of+Problem+Solving%22... From p.3 of the PDF: > Curating Cold Start RL Data: We constructed our initial training data through…

> Why do you think that the 2024 Putnam programs that they used to test were in the training data?

They reference https://artofproblemsolving.com/community/c13_contest_collec... for the source of their scrape and the Putnam problems are on that page under 'Undergraduate Contests'.

Re: DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]

#52

That is amazing if they can do all of this at < 10 % of the cost of frontier labs. Off course they work in the shadows of the great work done in the frontier labs and shared, but there is some exceptional high speed execution happening behind the scenes that shows this is clearly a race, but a race where China is happy to be #2 as long as the gap is not significant and the costs are reasonable

Frankly, I am pleasantly surprised to see that being a relatively close number two seems to be both practical and is turning out to be enormously beneficial to humanity. I am concerned that deep secrecy on OAIs part could change that, but it’s also possible that the genie is sufficiently out of the bottle that it no longer would be practical.

Re: DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]

#53

Amazing model! I'm trying to get it to run on an ec2 machine right now, but it looks like a lot of the performance actually depends on more than just classical LLM inference. And it looks like Deepseek didn't share their scripts to do the parallel thinking traces and self-verification loops. Is anybody else working on recreating this right now?

Hi! Did you ever end up running this reproduction? If yes, could you also check if the Putnam/IMO problems are in the training data perhaps by trying to have it complete the problems n times? I would totally do this myself if I weren’t GPU poor!
Post reply on HN