Tree Search Distillation for Language Models Using PPO
11–15 of 15 posts
I may never understand what harness means - it's used in so many contexts
Re: Tree Search Distillation for Language Models Using PPO
#12Why is almost every RL paper done on Qwen-2.5 ? That decreases its credibility.
It makes it easier to compare with other papers. If two different papers apply different methods to different models and get different results, how do you know which method is better?
Once you have identified the best method and want to productize it, it would of course make sense to apply it on top of the best model, but if you're just doing research, you can skip that expensive last step.
Re: Tree Search Distillation for Language Models Using PPO
#13Why is almost every RL paper done on Qwen-2.5 ? That decreases its credibility.
> Why is almost every RL paper done on Qwen-2.5 ?
In what way does using this model reduce the authors credibility?
Re: Tree Search Distillation for Language Models Using PPO
#14I may never understand what harness means - it's used in so many contexts
Its a thing that isn't part of the "subject", used with the subject, to manipulate the state of the "the subject" to be closer to what we want.
Re: Tree Search Distillation for Language Models Using PPO
#15Great post! I wonder why MCTS is not more popular as a test time compute harness. Did you compare performance of MCTS (without distillation) against other methods (eg best of N) with the same compute budget?
I didn't compare with the harness (focused on distillation) but the original ToT paper has a section on it: https://arxiv.org/abs/2305.10601