Live data from Hacker News

OpenAI Reinforcement Fine-Tuning Research Program

openai.com

1–10 of 67 posts

Re: OpenAI Reinforcement Fine-Tuning Research Program

#3

In a final lecture at UC Berkeley this semester, Dawn Song was very clear that malicious fine tuning is a top priority among implementers right now. "Towards building safe and trustworthy AI Agents and a Path for Science- and Evidence-based AI Policy."

Say more…

Re: OpenAI Reinforcement Fine-Tuning Research Program

#9
post #5

What are the advantages of reinforcement learning over DPO (Direct Preference Optimization)? My understanding is that the DPO paper showed it was equivalent to RLHF, but simpler and more computationally efficient.

you mean PPO not RLHF

simpler/efficient is not just about compute. its also data efficient.

Post reply on HN