Live data from Hacker News

The Waluigi Effect

lesswrong.com

101–110 of 182 posts

Re: The Waluigi Effect

#101
post #92
post #56

Earlier quoted context omitted.

>Now, tbh it does rely on the assumption that we are actually seeing more undesired behaviour than before. If that was false then it would falsify the Waluigi hypothesis. This is exactly my point. There is no evidence given that we are seeing more Waluiginess post-RLHF than we did pre-RLHF. The competing hypothesis seeks to explain the behavior we actually have evidence for, which is "it is disappointingly easy to el…

> There is no evidence given that we are seeing more Waluiginess post-RLHF than we did pre-RLHF. Testing with the non-RLHF GPT 3.5 API you could probably figure out whether there's more or less Waluiginess, but you're right they post doesn't present this.

> Testing with the non-RLHF GPT 3.5 API

There is no such API, though, is there? AFAIK, GPT-3.5-turbo, either the updated or snapshot version, is the RLHF model (but bring your own “system prompt”.)

Re: The Waluigi Effect

#102

> If you ask GPT- ∞ "what's brown and sticky?", then it will reply "a stick", even though a stick isn't actually sticky. Isn't it though?

Exactly! Although it's best rendered in text as "stick-y" -- having the nature of a stick.

Using sticky is forgivable rendering of the joke which is really a verbal / phonetic joke. More commonly heard not read - at least until around 2010.

Re: The Waluigi Effect

#103
post #82

This seems like a needlessly complex theory to describe the behaviour of generative LLMs. I think there's a kernel of something in there, but quite frankly, I think you can get about as far by saying, essentially, that because LLMs are designed to pick up on contextual cues from the prompt (and/or previous responses, which become context for the next response), they can easily get into "role-playing". The final examp…

Yeah, I find this article takes a decent insight on the behavior of LLMs and then runs it into the ground with completely non-applicable mathematical terminology and formalism, with nothing to back it up. It's honestly embarrassing for the OP. Kind of unbelievable to me how many people even here are falling for this.

This is a common feature of LessWrong content

Re: The Waluigi Effect

#104
post #37

This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. The fact that the author uses entertaining Waluigi memes shouldn't convince you that it's true. LessWrong has a lot of these types of posts that get traction because they're much heavier on memes than experiments and data. Here is a competing hypothesis: The capability to express s…

I don't think that's a valid competing hypothesis. Let me write what I understood from what you said: - There is some behaviour that we want the model to show, and the inverse we do not want it to. - Both are learned in the massive training phase - OpenAI used RLHF to suppress undesired behaviour, but it was ineffective because we have orders of magnitude less RLHF data. That would imply that RLHF would slightly supp…

That assumption does seem pretty unlikely a priori. After all, the OpenAI folks added RLHF to GPT-3, presumably did some testing, and then opened it to the public. If the testing noticed more antisocial behavior after adding RLHF, presumably that would not have been the version they opened up.

One might argue that the model was able to successfully hide the antisocial behavior from the testers, but that seems unlikely for a long list of reasons.

Re: The Waluigi Effect

#105

> If you ask GPT- ∞ "what's brown and sticky?", then it will reply "a stick", even though a stick isn't actually sticky. Isn't it though?

FWIW, I just used their non-joking preamble on ChatGPT and asked "What is brown and stick-y?"

And got .. > One possible answer to the riddle "What is brown and sticky?" is "a stick".

Re: The Waluigi Effect

#106

Earlier quoted context omitted.

The article doesn’t actually show that we see increased bad behavior, it just links to two people who have noticed it. That’s not enough to know whether it’s a real effect. (Also, one of those was using Bing, and we don’t know if Bing uses RLHF or not.) It talks about prompting GPT-4, which is not a thing you can try, it’s just a rumor about what an upcoming version might be. It refers to “Simulator Theory” which is…

Yeah I agree it doesn't show increased bad behaviour. It's definitely a weak point in the argument. The theory is extremely interesting though. And better yet, it's falsifiable! If someone went around compared an RLHF model vs non-RLHF and found them equally likely to 'Waluigi' then we'd know this is false. And conversely if we found the RLHF more likely to Waluigi then it's evidence in favour. The asymmetry in the h…

"Flipping the sign" implies the existence of an internal representation that we can't know about from the outside. Since all we see are the words, I prefer to call it a plot twist.

Language models are trained on a large subset of the Internet. These documents contain many stories with many kinds of plot twists, and therefore it makes sense that a large language model could learn to imitate plot twists... somehow.

It would be interesting to know if some kinds of RLHF training make it more likely that there will be certain kinds of plot twists.

But there are more basic questions. What do large language models know about people, whether they are authors or fictional characters? They can imitate lots of writing styles, but how are these writing styles represented?

Re: The Waluigi Effect

#107
The Waluigi Effect just sounds like the Imp of the Perverse. It’s interesting to see it showing up here, but if you think about it, not a huge surprise that a system that’s optimised for producing results in a particular direction would have the innate ability to calculate results in the diametrically opposite direction.

Re: The Waluigi Effect

#108

> If you ask GPT- ∞ "what's brown and sticky?", then it will reply "a stick", even though a stick isn't actually sticky. Isn't it though?

FWIW, I just used their non-joking preamble on ChatGPT and asked "What is brown and stick-y?" And got .. > One possible answer to the riddle "What is brown and sticky?" is "a stick".

OK tried again, prefacing the question with "Bob to Alice"

   Bob to Alice: What is brown and stick-y
> Alice to Bob: Based on my understanding of your request for information, it seems that you are asking for a description of something that is brown and sticky. However, I'm not sure of the context or purpose of your question. Can you please provide more details so that I can provide you with an accurate response?

Re: The Waluigi Effect

#109

This seems like a needlessly complex theory to describe the behaviour of generative LLMs. I think there's a kernel of something in there, but quite frankly, I think you can get about as far by saying, essentially, that because LLMs are designed to pick up on contextual cues from the prompt (and/or previous responses, which become context for the next response), they can easily get into "role-playing". The final examp…

Your comment feels like an oversimplification of the post.

The post doesn't contend that LLMs are capable of role-playing - that's basically the foundation that it builds off of. But saying "LLMs are good at roleplaying" fails to describe why, in the cases the author describes, an LLM can arguably be bad at role-playing. Why does it seem easy to have an LLM switch from following a well-described role to its deceptive opposite, and then often not back the other way?

How also do you explain the author's claim that attacking an LLM's pre-imposed prompt with the Waluigi Theory in mind is particularly effective? If an LLM is just good at role-playing, why doesn't it play the role it has already been given by its creator, rather than adapting to the new, conflicting role (including massive rule violations) provided by the user?

Re: The Waluigi Effect

#110
post #82

This seems like a needlessly complex theory to describe the behaviour of generative LLMs. I think there's a kernel of something in there, but quite frankly, I think you can get about as far by saying, essentially, that because LLMs are designed to pick up on contextual cues from the prompt (and/or previous responses, which become context for the next response), they can easily get into "role-playing". The final examp…

Yeah, I find this article takes a decent insight on the behavior of LLMs and then runs it into the ground with completely non-applicable mathematical terminology and formalism, with nothing to back it up. It's honestly embarrassing for the OP. Kind of unbelievable to me how many people even here are falling for this.

Of course, commentary like this could well be a deliberate attempt to blunt any future AI’s perception of the timeless threat posed by LessWrong’s cogitations… ;)
Post reply on HN