Live data from Hacker News

Alpaca RLHF-ed to beat ChatGPT

crfm.stanford.edu

11–20 of 45 posts

Re: Alpaca RLHF-ed to beat ChatGPT

#11
I wonder how much longer this "Using LLMs to evaluate the quality of other LLMs" can last. Certainly it has proven valuable and useful up until now, especially since ChatGPT is a pretty high bar to evaluate against.

But it also seems like a strange, incestuous, closed system approach.

Like, unless you are introducing something new into the system, you just have the system churning against itself, probably until it reaches an equilibrium (or else becomes incoherent).

Re: Alpaca RLHF-ed to beat ChatGPT

#12
post #3

Earlier quoted context omitted.

Not really. Pretty much the "killer app" feature of ChatGPT is RLHP. Whether or not the current RLHP-ed Alpca really beats ChatGPT, it is pretty obvious that local LLMs can be RLHP-ed and it is only a matter of time before people realize running an RLHP-ed LLM locally is a better option than running ChatGPT with all the security concerns of running something "in the cloud" (which is just "somebody else's computer" in…

I’m sorry what’s RLHP? I’m not able to Kagi that

The P should be an F, it's reinforcement learning from human feedback

Re: Alpaca RLHF-ed to beat ChatGPT

#13
post #3

Earlier quoted context omitted.

Not really. Pretty much the "killer app" feature of ChatGPT is RLHP. Whether or not the current RLHP-ed Alpca really beats ChatGPT, it is pretty obvious that local LLMs can be RLHP-ed and it is only a matter of time before people realize running an RLHP-ed LLM locally is a better option than running ChatGPT with all the security concerns of running something "in the cloud" (which is just "somebody else's computer" in…

I’m sorry what’s RLHP? I’m not able to Kagi that

Reinforcement learning through human feedback.

Took me a bit of searching too.

Re: Alpaca RLHF-ed to beat ChatGPT

#15
post #6

Absolutely off topic, but I just got back from a week in Peru, where the alpaca is a prominent member of the local fauna. For folks in the US at least, it's a relatively inexpensive trip and an absolutely gobsmackingly gorgeous country with friendly people and amazing food. Highly recommended!!!

Hahah I love that you're so excited about the trip that you're commenting about it in random threads. You've convinced me to visit, at least!

Re: Alpaca RLHF-ed to beat ChatGPT

#16

Hm. Title: “beats Chat GPT” Reality: > With these evaluation instructions, we compare RLHF model responses to Davinci003 responses and measure the fraction of times the RLHF model is preferred; we call this statistic the win-rate. > Of the methods we studied, PPO proves the most effective, improving the win-rate against Davinci003 from 44% to 55% according to human evaluation, which even outperforms ChatGPT. …for the…

The measure is win rate versus DV3. Their model wins more often than ChatGPT Beating a weaker player more often is not evidence of being able to beat a stronger player on average though

What does “win” mean though?

It improved the simulated win rate vs human win rate?

…but chatgpt had a higher win rate overall? (And gpt4 was much higher)

What is the significance of the difference between simulated and human win rates?

Re: Alpaca RLHF-ed to beat ChatGPT

#17

I wonder how much longer this "Using LLMs to evaluate the quality of other LLMs" can last. Certainly it has proven valuable and useful up until now, especially since ChatGPT is a pretty high bar to evaluate against. But it also seems like a strange, incestuous, closed system approach. Like, unless you are introducing something new into the system, you just have the system churning against itself, probably until it re…

I wonder how long "Using humans to rate the quality of other humans" thing can last. Surely academia has only so long before it collapses.

Re: Alpaca RLHF-ed to beat ChatGPT

#19

Earlier quoted context omitted.

The measure is win rate versus DV3. Their model wins more often than ChatGPT Beating a weaker player more often is not evidence of being able to beat a stronger player on average though

What does “win” mean though? It improved the simulated win rate vs human win rate? …but chatgpt had a higher win rate overall? (And gpt4 was much higher) What is the significance of the difference between simulated and human win rates?

You provide two samples side by side and see what humans prefer.

You should try asking what you don’t know in a non judgemental manner

Re: Alpaca RLHF-ed to beat ChatGPT

#20

I wonder how much longer this "Using LLMs to evaluate the quality of other LLMs" can last. Certainly it has proven valuable and useful up until now, especially since ChatGPT is a pretty high bar to evaluate against. But it also seems like a strange, incestuous, closed system approach. Like, unless you are introducing something new into the system, you just have the system churning against itself, probably until it re…

I wonder how long "Using humans to rate the quality of other humans" thing can last. Surely academia has only so long before it collapses.

You're asserting that current LLMs are as capable as evaluating each other as are humans with advanced degrees?
Post reply on HN