Something weird here, why is it so hard to have a deterministic program capable of checking a proof or anything math related, aren't maths super deterministic when natural language is not. From first principles, it should be possible to do this without a llm verifier.
DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]
11–20 of 53 posts
Re: DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]
#12Re: DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]
#13We've seen absolutely ridiculous progress in model capability over the past year (which is also quite terrifying).
Re: DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]
#14Re: DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]
#15It's cool, but I genuinely cannot fathom why they are targeting natural language proofs instead of a proof assistant.
But I suppose the bigger goal remains improving their language model, and this was an experimentation born from that. These works are symbiotic; the original DeepSeekMath resulted in GRPO, which eventually formed the backbone of their R1 model: https://arxiv.org/abs/2402.03300
Re: DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]
#16It's cool, but I genuinely cannot fathom why they are targeting natural language proofs instead of a proof assistant.
Re: DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]
#17Something weird here, why is it so hard to have a deterministic program capable of checking a proof or anything math related, aren't maths super deterministic when natural language is not. From first principles, it should be possible to do this without a llm verifier.
Plus there isn't a lot of training data in lean.
Most gains come from training on stuff already out there, not really the RLVR part which just amps it up a bit.
Re: DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]
#18So it's designed for informal proofs and it "verifies" based on a rubric fitting function and human interaction, is that right? What's the use case for a system like this?
I suspect it's also because there isn't a lot of data to train on.
Re: DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]
#19If i read it right it used multiple samples of itself to verify the aqccuracy, but isnt this problematic?
Re: DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning [pdf]
#20If i read it right it used multiple samples of itself to verify the aqccuracy, but isnt this problematic?