OpenAI Reinforcement Fine-Tuning Research Program
1–10 of 67 posts
Re: OpenAI Reinforcement Fine-Tuning Research Program
#2In a final lecture at UC Berkeley this semester, Dawn Song was very clear that malicious fine tuning is a top priority among implementers right now.
"Towards building safe and trustworthy AI Agents and a Path for Science- and Evidence-based AI Policy."
Re: OpenAI Reinforcement Fine-Tuning Research Program
#3In a final lecture at UC Berkeley this semester, Dawn Song was very clear that malicious fine tuning is a top priority among implementers right now. "Towards building safe and trustworthy AI Agents and a Path for Science- and Evidence-based AI Policy."
Say more…
Re: OpenAI Reinforcement Fine-Tuning Research Program
#4Who owns the fine tuning IP. Can OpenAI resell your model after investing a lot in it?
Re: OpenAI Reinforcement Fine-Tuning Research Program
#5What are the advantages of reinforcement learning over DPO (Direct Preference Optimization)? My understanding is that the DPO paper showed it was equivalent to RLHF, but simpler and more computationally efficient.
Re: OpenAI Reinforcement Fine-Tuning Research Program
#6This was announced as part of their second day of "12 Days of AI": https://www.youtube.com/watch?v=fMJMhBFa_Gc
Re: OpenAI Reinforcement Fine-Tuning Research Program
#7Clever way to get more training data.
Re: OpenAI Reinforcement Fine-Tuning Research Program
#8this sounds like expert systems 2.0 lol
Re: OpenAI Reinforcement Fine-Tuning Research Program
#9What are the advantages of reinforcement learning over DPO (Direct Preference Optimization)? My understanding is that the DPO paper showed it was equivalent to RLHF, but simpler and more computationally efficient.
you mean PPO not RLHF
simpler/efficient is not just about compute. its also data efficient.
Re: OpenAI Reinforcement Fine-Tuning Research Program
#10Clever way to get more training data.
Yeah I was gonna say this would normally be paid for. They're profiting off of the hype.