LIMO: Less Is More for Reasoning
11–20 of 137 posts
Re: LIMO: Less Is More for Reasoning
#12I think I've recently read two seemingly contradicting things: 1- LLMs can never generalize theorem proving 2- this paper: "This suggests that contemporary LLMs may already possess rich mathematical knowledge in their parameter space, transforming the challenge from knowledge acquisition to knowledge elicitation" Not sure what is what anymore!
Re: LIMO: Less Is More for Reasoning
#13where's chatbotAI-zero? in the way alpha-go-zero was the best after training with itself? (and only with itself)
The advantage of alpha-go-zero is that it is constrained to the language of go. If you made two LLM train only off each other they would develop their own language. Maybe they'd be great at reasoning, but we wouldn't understand them. Even humans in that situation would develop jargon, and as time goes on a dialect or language of their own. And humans are a lot more grounded in their language than LLMs.
Re: LIMO: Less Is More for Reasoning
#14In the same way that image diffusion models showed that convincing approximations of the entire visual world could be summarized in a 5GB model, are "reasoning patterns" similarly compressible? Are there actually countably few reasoning patterns that are used across all domains, and as such can be captured with relatively small training sets?
Perhaps in a domain like math a smallish number of math-specific reasoning steps will go a long way, but math itself also has many "sub-domains" (algebra, geometry, calculus, topology, etc) and AFAIK the techniques of one branch are only going to be useful in another to extent you can map the problem from one domain to another.
Re: LIMO: Less Is More for Reasoning
#15where's chatbotAI-zero? in the way alpha-go-zero was the best after training with itself? (and only with itself)
Thankfully for mathematics and code this seems plausible due to automated theorem proving.
Re: LIMO: Less Is More for Reasoning
#16where's chatbotAI-zero? in the way alpha-go-zero was the best after training with itself? (and only with itself)
> With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing.
A lot of these we can probably solve, but as other have pointed out we want a model that humans can converse with, not an AI for the purpose of other AI.
That said, it seems like a promising area of research:
> DeepSeek-R1-Zero demonstrates capabilities such as self-verification, reflection, and generating long CoTs, marking a significant milestone for the research community.
Re: LIMO: Less Is More for Reasoning
#17where's chatbotAI-zero? in the way alpha-go-zero was the best after training with itself? (and only with itself)
You don't want that as a product, in the sense that having an AI model train itself by simply having internal conversations without ever looking at any human-written content, might result in something that humans cannot comprehend. Also, well - there's the technicality of "you don't 'win' a conversation like you can 'win' at Go", so how would you know to reward the model as you're training it?
Re: LIMO: Less Is More for Reasoning
#18Re: LIMO: Less Is More for Reasoning
#19Re: LIMO: Less Is More for Reasoning
#20To see a World in a Grain of Sand And a Heaven in a Wild Flower, Hold Infinity in the palm of your hand And Eternity in an hour.