Live data from Hacker News

Alpaca RLHF-ed to beat ChatGPT

crfm.stanford.edu

1–10 of 45 posts

Re: Alpaca RLHF-ed to beat ChatGPT

#3
post #2

The title is a bit much, no?

Not really. Pretty much the "killer app" feature of ChatGPT is RLHP. Whether or not the current RLHP-ed Alpca really beats ChatGPT, it is pretty obvious that local LLMs can be RLHP-ed and it is only a matter of time before people realize running an RLHP-ed LLM locally is a better option than running ChatGPT with all the security concerns of running something "in the cloud" (which is just "somebody else's computer" in the famous saying).

Re: Alpaca RLHF-ed to beat ChatGPT

#5
post #3
post #2

The title is a bit much, no?

Not really. Pretty much the "killer app" feature of ChatGPT is RLHP. Whether or not the current RLHP-ed Alpca really beats ChatGPT, it is pretty obvious that local LLMs can be RLHP-ed and it is only a matter of time before people realize running an RLHP-ed LLM locally is a better option than running ChatGPT with all the security concerns of running something "in the cloud" (which is just "somebody else's computer" in…

I was referring to the HN guidelines against editorializing titles.

Re: Alpaca RLHF-ed to beat ChatGPT

#6
Absolutely off topic, but I just got back from a week in Peru, where the alpaca is a prominent member of the local fauna.

For folks in the US at least, it's a relatively inexpensive trip and an absolutely gobsmackingly gorgeous country with friendly people and amazing food. Highly recommended!!!

Re: Alpaca RLHF-ed to beat ChatGPT

#7
Hm. Title: “beats Chat GPT”

Reality:

> With these evaluation instructions, we compare RLHF model responses to Davinci003 responses and measure the fraction of times the RLHF model is preferred; we call this statistic the win-rate.

> Of the methods we studied, PPO proves the most effective, improving the win-rate against Davinci003 from 44% to 55% according to human evaluation, which even outperforms ChatGPT.

…for the metric we invented, which measures… the difference between a simulated and human evaluated result.

Or something.

Does anyone have a good idea of what this metric actually means and if it is actually relevant to anything useful?

Re: Alpaca RLHF-ed to beat ChatGPT

#8
post #3
post #2

The title is a bit much, no?

Not really. Pretty much the "killer app" feature of ChatGPT is RLHP. Whether or not the current RLHP-ed Alpca really beats ChatGPT, it is pretty obvious that local LLMs can be RLHP-ed and it is only a matter of time before people realize running an RLHP-ed LLM locally is a better option than running ChatGPT with all the security concerns of running something "in the cloud" (which is just "somebody else's computer" in…

I’m sorry what’s RLHP? I’m not able to Kagi that

Re: Alpaca RLHF-ed to beat ChatGPT

#9
post #6

Absolutely off topic, but I just got back from a week in Peru, where the alpaca is a prominent member of the local fauna. For folks in the US at least, it's a relatively inexpensive trip and an absolutely gobsmackingly gorgeous country with friendly people and amazing food. Highly recommended!!!

[flagged]

Re: Alpaca RLHF-ed to beat ChatGPT

#10

Hm. Title: “beats Chat GPT” Reality: > With these evaluation instructions, we compare RLHF model responses to Davinci003 responses and measure the fraction of times the RLHF model is preferred; we call this statistic the win-rate. > Of the methods we studied, PPO proves the most effective, improving the win-rate against Davinci003 from 44% to 55% according to human evaluation, which even outperforms ChatGPT. …for the…

The measure is win rate versus DV3. Their model wins more often than ChatGPT

Beating a weaker player more often is not evidence of being able to beat a stronger player on average though

Post reply on HN