Live data from Hacker News

The Waluigi Effect

lesswrong.com

91–100 of 182 posts

Re: The Waluigi Effect

#91

I find it fascinating that AI alarmists spent years writing gigabytes of text scaring themselves about how an unaligned AI would behave, and are now feeding that into training models that teach a pretty capable AI how to act. We've talked in the past about how transhumanism is a religion that creates its own God, but this is an even funnier example where vastly intelligent people are optimizing a software system to s…

Go deeper - they are now writing text about how writing text about rogue AIs might create a rogue AI...

https://gwern.net/fiction/clippy

Re: The Waluigi Effect

#92
post #56

Earlier quoted context omitted.

I don't think that's a valid competing hypothesis. Let me write what I understood from what you said: - There is some behaviour that we want the model to show, and the inverse we do not want it to. - Both are learned in the massive training phase - OpenAI used RLHF to suppress undesired behaviour, but it was ineffective because we have orders of magnitude less RLHF data. That would imply that RLHF would slightly supp…

>Now, tbh it does rely on the assumption that we are actually seeing more undesired behaviour than before. If that was false then it would falsify the Waluigi hypothesis. This is exactly my point. There is no evidence given that we are seeing more Waluiginess post-RLHF than we did pre-RLHF. The competing hypothesis seeks to explain the behavior we actually have evidence for, which is "it is disappointingly easy to el…

> There is no evidence given that we are seeing more Waluiginess post-RLHF than we did pre-RLHF.

Testing with the non-RLHF GPT 3.5 API you could probably figure out whether there's more or less Waluiginess, but you're right they post doesn't present this.

Re: The Waluigi Effect

#93

Earlier quoted context omitted.

I don't think that's a valid competing hypothesis. Let me write what I understood from what you said: - There is some behaviour that we want the model to show, and the inverse we do not want it to. - Both are learned in the massive training phase - OpenAI used RLHF to suppress undesired behaviour, but it was ineffective because we have orders of magnitude less RLHF data. That would imply that RLHF would slightly supp…

The article doesn’t actually show that we see increased bad behavior, it just links to two people who have noticed it. That’s not enough to know whether it’s a real effect. (Also, one of those was using Bing, and we don’t know if Bing uses RLHF or not.) It talks about prompting GPT-4, which is not a thing you can try, it’s just a rumor about what an upcoming version might be. It refers to “Simulator Theory” which is…

Yeah I agree it doesn't show increased bad behaviour. It's definitely a weak point in the argument.

The theory is extremely interesting though. And better yet, it's falsifiable! If someone went around compared an RLHF model vs non-RLHF and found them equally likely to 'Waluigi' then we'd know this is false. And conversely if we found the RLHF more likely to Waluigi then it's evidence in favour.

The asymmetry in the hypothesis is really nice too. If this was true then I'd expect it to be possible to flip the sign in the RLHF step, effectively training it in favour of 'bad' behaviour. Then forcefully inducing 'Waluigi collapse' before opening to the public!

Re: The Waluigi Effect

#94
post #80
post #76

Earlier quoted context omitted.

> What's actually going on is that a LLM is like the language center of a brain, without the brain. I’ve seen this sentiment expressed multiple times, but is that really correct? Maybe this works differently for other people, but I’ve noticed that I have to use my language to really think. I can do trivial things mindlessly, but to solve a problem, I need to express it with words in my mind. It makes me feel like the…

I know I am in the minority out there but when I do math, calculus, diff eq, whatever, the answer just comes to me. There's no internal dialogue, the answer just, for the lack of a better phrase, rises from the deep and is known to me. When I am in an discussion I will look up and off the my left when I am thinking, but no words are happening in my "inner dialogue", it's just nothing and then I start speaking whateve…

I think that math and programming may be the odd exceptions, as they employ their own language-like constructs. There might be a misconception in the name of LLMs - we say that they’re language models, but really they’re token models, some of which may be human language words, while others may represent other things.

As for the discussions, I agree that I don’t have a distinct narrative in my mind during one, but I also noticed that I don’t really know what exactly I’m going to say when I start a response. So it also feels like the act of responding is actually heavily involved in creating the response, rather than just putting it into words.

BTW, I’ve always wondered if people really think differently, or we just describe it in different ways. I guess we’ll never really know.

Re: The Waluigi Effect

#95
While the formality is way overwrought (and ChatGPT is not creating any "simulacra" of characters), I think the overall point is correct that language models are trained on stories and other human writing, and inversion is a very common plot point in stories and human writing in general, if only because contradicting expectations is more interesting (e.g. "man bites dog").

We also less commonly see exposition that is not germane to a story, so a character is rarely even mentioned to be "weak", "intelligent", etc unless there is a point. And sometimes the point is that they are later shown to be "strong", "absent-minded", or other contradictions. Which means that mentioning a character's strength makes it more likely they will later be described as weak, than if it was never mentioned at all. Finally, double-contradiction is less common in human text (maybe because plain contradiction is sufficiently interesting), so a running text with no reversals is more likely to eventually reverse, than a running text with one reversal is to return to its original state.

While I don't agree at all with the author's sense that this represents some kind of "alignment" danger, it does go a long way to explaining why ChatGPT is easy to pull into conversations that shock or surprise, despite all the training. It's because human writing often attempts to shock and surprise, and the LLM is training on that statistically.

Re: The Waluigi Effect

#96
post #37

This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. The fact that the author uses entertaining Waluigi memes shouldn't convince you that it's true. LessWrong has a lot of these types of posts that get traction because they're much heavier on memes than experiments and data. Here is a competing hypothesis: The capability to express s…

>This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. Applicable to much of the rationalist AI risk discourse.

I hesitate to defend AI safety discourse, but I will say that philosophy in general is sort of fanficy, and AI safety is something I'd loosely associate with philosophy.

Re: The Waluigi Effect

#97

Does the scientific community at large take the theories of these LessWrong-type "researchers" seriously? Sounds like a bunch of mumbo jumbo to me, with some LaTeX sprinkled in to look more serious.

No, and it's a huge sticking point that the AI safety group is super salty about. They call themselves scientists and researchers and get super defensive when actual researchers (people who have PhDs and get published in journals) imply that they aren't.

Re: The Waluigi Effect

#98

I find it fascinating that AI alarmists spent years writing gigabytes of text scaring themselves about how an unaligned AI would behave, and are now feeding that into training models that teach a pretty capable AI how to act. We've talked in the past about how transhumanism is a religion that creates its own God, but this is an even funnier example where vastly intelligent people are optimizing a software system to s…

I can't imagine a LLM trained on the entirety of the internet would have any material influence from writings around AI safety

Re: The Waluigi Effect

#99
post #82

This seems like a needlessly complex theory to describe the behaviour of generative LLMs. I think there's a kernel of something in there, but quite frankly, I think you can get about as far by saying, essentially, that because LLMs are designed to pick up on contextual cues from the prompt (and/or previous responses, which become context for the next response), they can easily get into "role-playing". The final examp…

Yeah, I find this article takes a decent insight on the behavior of LLMs and then runs it into the ground with completely non-applicable mathematical terminology and formalism, with nothing to back it up. It's honestly embarrassing for the OP. Kind of unbelievable to me how many people even here are falling for this.

Sounds about right for the increasingly ironically-named LessWrong site…

Re: The Waluigi Effect

#100
post #96

Earlier quoted context omitted.

>This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. Applicable to much of the rationalist AI risk discourse.

I hesitate to defend AI safety discourse, but I will say that philosophy in general is sort of fanficy, and AI safety is something I'd loosely associate with philosophy.

> > Applicable to much of the rationalist AI risk discourse

> I hesitate to defend AI safety discourse

The rationalist AI risk discourse is not the same thing as AI safety discourse, in any case; it’s a small corner of the larger whole.

Post reply on HN