Alpaca RLHF-ed to beat ChatGPT
crfm.stanford.edu
Alpaca RLHF-ed to beat ChatGPT
1–10 of 45 posts
Re: Alpaca RLHF-ed to beat ChatGPT
#2Re: Alpaca RLHF-ed to beat ChatGPT
#3The title is a bit much, no?
Re: Alpaca RLHF-ed to beat ChatGPT
#4The title is a bit much, no?
Re: Alpaca RLHF-ed to beat ChatGPT
#5The title is a bit much, no?
Not really. Pretty much the "killer app" feature of ChatGPT is RLHP. Whether or not the current RLHP-ed Alpca really beats ChatGPT, it is pretty obvious that local LLMs can be RLHP-ed and it is only a matter of time before people realize running an RLHP-ed LLM locally is a better option than running ChatGPT with all the security concerns of running something "in the cloud" (which is just "somebody else's computer" in…
Re: Alpaca RLHF-ed to beat ChatGPT
#6For folks in the US at least, it's a relatively inexpensive trip and an absolutely gobsmackingly gorgeous country with friendly people and amazing food. Highly recommended!!!
Re: Alpaca RLHF-ed to beat ChatGPT
#7Reality:
> With these evaluation instructions, we compare RLHF model responses to Davinci003 responses and measure the fraction of times the RLHF model is preferred; we call this statistic the win-rate.
> Of the methods we studied, PPO proves the most effective, improving the win-rate against Davinci003 from 44% to 55% according to human evaluation, which even outperforms ChatGPT.
…for the metric we invented, which measures… the difference between a simulated and human evaluated result.
Or something.
Does anyone have a good idea of what this metric actually means and if it is actually relevant to anything useful?
Re: Alpaca RLHF-ed to beat ChatGPT
#8The title is a bit much, no?
Not really. Pretty much the "killer app" feature of ChatGPT is RLHP. Whether or not the current RLHP-ed Alpca really beats ChatGPT, it is pretty obvious that local LLMs can be RLHP-ed and it is only a matter of time before people realize running an RLHP-ed LLM locally is a better option than running ChatGPT with all the security concerns of running something "in the cloud" (which is just "somebody else's computer" in…
Re: Alpaca RLHF-ed to beat ChatGPT
#9Absolutely off topic, but I just got back from a week in Peru, where the alpaca is a prominent member of the local fauna. For folks in the US at least, it's a relatively inexpensive trip and an absolutely gobsmackingly gorgeous country with friendly people and amazing food. Highly recommended!!!
Re: Alpaca RLHF-ed to beat ChatGPT
#10Hm. Title: “beats Chat GPT” Reality: > With these evaluation instructions, we compare RLHF model responses to Davinci003 responses and measure the fraction of times the RLHF model is preferred; we call this statistic the win-rate. > Of the methods we studied, PPO proves the most effective, improving the win-rate against Davinci003 from 44% to 55% according to human evaluation, which even outperforms ChatGPT. …for the…
Beating a weaker player more often is not evidence of being able to beat a stronger player on average though