Don't Build an RL Environment Startup
benanderson.work
Don't Build an RL Environment Startup
1–2 of 2 posts
Re: Don't Build an RL Environment Startup
#2Reinforcement Learning from Human Feedback [1] will be published next week. Hard to believe the technique is already obsolete.
[1] https://www.manning.com/books/reinforcement-learning-from-hu...