ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
11–20 of 34 posts
Re: ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
#12Cool cool. I'm a bit put off by calling it "reasoning" /"thought". These RL targets can be achieved without "thinking" model but still cool. Gotta love the brainfuck task. I personally think that Gemini 2.5 Pro's superiority comes from having hundreds or thousands RL tasks (without any proof whatsoever, so rather a feeling). So I've been wanting a "RL Zoo" for quite a while. I hope this project won't be a one-off and…
Re: ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
#13Cool cool. I'm a bit put off by calling it "reasoning" /"thought". These RL targets can be achieved without "thinking" model but still cool. Gotta love the brainfuck task. I personally think that Gemini 2.5 Pro's superiority comes from having hundreds or thousands RL tasks (without any proof whatsoever, so rather a feeling). So I've been wanting a "RL Zoo" for quite a while. I hope this project won't be a one-off and…
We definitely plan to maintain the project for as long as there is interest in it. If you have ideas for new tasks, we'd always welcome contributions!
Re: ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
#14Re: ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
#15Re: ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
#16Cool to see NVIDIA’s most recent reasoning model [1] already uses Reasoning Gymas a large part of their data mixture [1] https://arxiv.org/abs/2505.24864
> prolonged RL training can uncover novel reasoning strategies that are inaccessible to base models, even under extensive sampling does this mean that previous RL papers claiming the opposite were possibly bottlenecked by small datasets?
Re: ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
#17Cool cool. I'm a bit put off by calling it "reasoning" /"thought". These RL targets can be achieved without "thinking" model but still cool. Gotta love the brainfuck task. I personally think that Gemini 2.5 Pro's superiority comes from having hundreds or thousands RL tasks (without any proof whatsoever, so rather a feeling). So I've been wanting a "RL Zoo" for quite a while. I hope this project won't be a one-off and…
Gemini 2.5 Pro's superiority is IMO largely driven by their long context support and training methodology. Compare Gemini as a beta reader for a 100k token book with GPT4.1 or Claude 4, and it becomes quite clear how much more effectively it can reason across its context than other comparable models. This also makes it much better for architecting new features into a system, since you can load a lot of the current sy…
Re: ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
#18[flagged]
Re: ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
#19by the love of god, please stop overfitting on gsm8k
Prejudices is a form of overfitting IMHO
Re: ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
#20by the love of god, please stop overfitting on gsm8k