Recent reasoning research: GRPO tweaks, base model RL and data curation #1 Post by ydnyshhh » Mon, Mar 31, 2025, 4:53 PM UTC Recent reasoning research: GRPO tweaks, base model RL and data curationinterconnects.ai