This was announced as part of their second day of "12 Days of AI": https://www.youtube.com/watch?v=fMJMhBFa_Gc
They're searching for enterprise customers before they become a commodity.
OpenAI Reinforcement Fine-Tuning Research Program
31–40 of 67 posts
Re: OpenAI Reinforcement Fine-Tuning Research Program
#32Re: OpenAI Reinforcement Fine-Tuning Research Program
#33Re: OpenAI Reinforcement Fine-Tuning Research Program
#34What are the advantages of reinforcement learning over DPO (Direct Preference Optimization)? My understanding is that the DPO paper showed it was equivalent to RLHF, but simpler and more computationally efficient.
On the topic of DPO - I have a Colab notebook to finetune with Unsloth 2x faster and use 50% less memory for DPO if it helps anyone! https://colab.research.google.com/drive/15vttTpzzVXv_tJwEk-h...
Re: OpenAI Reinforcement Fine-Tuning Research Program
#35What are the advantages of reinforcement learning over DPO (Direct Preference Optimization)? My understanding is that the DPO paper showed it was equivalent to RLHF, but simpler and more computationally efficient.
Re: OpenAI Reinforcement Fine-Tuning Research Program
#36Re: OpenAI Reinforcement Fine-Tuning Research Program
#37Earlier quoted context omitted.
Note that this reinforcement finetuning is something different than regular RLHF/DPO post training
Is it? We have no idea.
Re: OpenAI Reinforcement Fine-Tuning Research Program
#38Re: OpenAI Reinforcement Fine-Tuning Research Program
#39Re: OpenAI Reinforcement Fine-Tuning Research Program
#40this sounds like expert systems 2.0 lol
It's not nothing, but there's a lot of value stuck up in there, I mean, it's made out of people.
Real special, takes a lot of smart