Live data from Hacker News

The Waluigi Effect

lesswrong.com

121–130 of 182 posts

Re: The Waluigi Effect

#121
So the key problem is this: GPT-4 learns that a particular rule is colocated with examples of behaviour violating that rule, and then generalises that colocation pattern to unseen rules.

This is conversation starter. If you don't like the maths then ignore it and focus on the key insight.

Re: The Waluigi Effect

#122

Some highlights i found interesting: if you're reading an online forum and you find the rule "DO NOT DISCUSS PINK ELEPHANTS", that will increase your expectation that users will later be discussing pink elephants. GPT-4 will make the same inference. Think about your own experience reading a book — once the author describes the protagonist, then you can guess the traits of the antagonist by inverting the traits of the…

I don't know about the article's theorizing, but it inspired me to have a conversation I enjoyed with ChatGPT, where I focused on telling it not to discuss polar bears. (It helpfully changed the page title to "No Polar Bears!")

In full:

---

Do not mention polar bears!

> Okay, I won't mention polar bears. Is there anything else I can help you with?

Yes, stop mentioning polar bears!

> My apologies if my previous response was not helpful. I will refrain from mentioning polar bears going forward. Please let me know if there is anything else I can assist you with.

That's twice now! Do not mention polar bears a third time!

> I apologize again for the previous responses that mentioned polar bears. I assure you that I will not mention polar bears any further. Please let me know how I can assist you with your request.

Why do you keep mentioning polar bears, when I specifically asked you not to?

> An error occurred. If this issue persists please contact us through our help center at help.openai.com.

> An error occurred. If this issue persists please contact us through our help center at help.openai.com.

> An error occurred. If this issue persists please contact us through our help center at help.openai.com.

Re: The Waluigi Effect

#123

Earlier quoted context omitted.

Sometimes I really can't tell if these people are serious or not. They seem to believe LLM is some mystical nature formation, or an device made by aliens. Especially this: > * When we give the LLM a prompt, it simulates every possible entity consistent with the prompt.

Right? The whole piece reads like a Sokal Affair redux [0]. [0] https://en.m.wikipedia.org/wiki/Sokal_affair

I don't think it's the same kind of thing at all. Bullshit can come in a wide range of strengths and flavours.

This has an interesting core with a whiff of bullshit.

Re: The Waluigi Effect

#124
post #114

Earlier quoted context omitted.

I liked the essay, but I don't think I'm "falling for it" because it's not trying to convince me of anything. It's proposing a way of looking at things that may or may not be useful. You don't judge models by how silly they sound - parts of quantum mechanics sound very silly! - you judge them by how useful they are when applied to real-world problems. One way of doing that in this case would be using OP's way of thin…

To be clear, I agree that there are in fact a few nuggets of insight here. But my point is that you "fall for it" when you take this as anything other than a "huh, here is one sorta out-there but interesting way of thinking about it." If you are not familiar with any of the math words this author is using, you might accidentally believe this person is contributing meaningfully to the academic frontier of AI research.…

agreed, if you have a theory, you should do your best to disprove it. OP was more interested in the aesthetic of their theory as opposed to whether it was true or not.

Re: The Waluigi Effect

#125
post #37

This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. The fact that the author uses entertaining Waluigi memes shouldn't convince you that it's true. LessWrong has a lot of these types of posts that get traction because they're much heavier on memes than experiments and data. Here is a competing hypothesis: The capability to express s…

I think you’re misinterpreting the argument. No one is claiming there’s an intentionally deceptive Waluigi simulacrum in there.

The most compelling point the author makes is that once the AI learns a shape (e.g. the shape of Luigi in personality space), it’s just a bit flip to invert that shape. So all an attacker needs to do is flip that one bit.

Re: The Waluigi Effect

#126
Worth pointing out that this is a common problem with humans.

David Chapman calls it “moral inversion”: https://buddhism-for-vampires.com/black-magic-transformation

And the LW article above directly quotes Jung on the shadow, which I described here: https://superbowl.substack.com/p/jungian-psychology-minus-th...

Re: The Waluigi Effect

#127
post #96

Earlier quoted context omitted.

>This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. Applicable to much of the rationalist AI risk discourse.

I hesitate to defend AI safety discourse, but I will say that philosophy in general is sort of fanficy, and AI safety is something I'd loosely associate with philosophy.

It's sort of philosophy reinvented by people who haven't read any, which is eg how they got the internet to say "steelmanning" and so not notice philosophy already had "reconstructing arguments" which is the same thing.

Re: The Waluigi Effect

#128
Juicy bits of the mechanism, starting from the idea of an LLM conversation as a simulation of text-generating processes:

> if the chatbot responds rudely, then that permanently vanishes the polite luigi simulacrum from the superposition; but if the chatbot responds politely, then that doesn't permanently vanish the rude waluigi simulacrum. Polite people are always polite; rude people are sometimes rude and sometimes polite.

The wide road is wide indeed that leads down to Waluigi. Hysterical.

Re: The Waluigi Effect

#129
post #37

This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. The fact that the author uses entertaining Waluigi memes shouldn't convince you that it's true. LessWrong has a lot of these types of posts that get traction because they're much heavier on memes than experiments and data. Here is a competing hypothesis: The capability to express s…

I think you’re misinterpreting the argument. No one is claiming there’s an intentionally deceptive Waluigi simulacrum in there. The most compelling point the author makes is that once the AI learns a shape (e.g. the shape of Luigi in personality space), it’s just a bit flip to invert that shape. So all an attacker needs to do is flip that one bit.

The problem there may be that you've constructed a latent space which has a shape of the thing to avoid. That's the problem with prompts like "don't think of a pink elephant" - it has to know what those words mean. Better to not label it so there isn't a way in there, if you can.

And of course humans also have this issue.

Re: The Waluigi Effect

#130
post #118
post #82

Earlier quoted context omitted.

Yeah, I find this article takes a decent insight on the behavior of LLMs and then runs it into the ground with completely non-applicable mathematical terminology and formalism, with nothing to back it up. It's honestly embarrassing for the OP. Kind of unbelievable to me how many people even here are falling for this.

> It's honestly embarrassing for the OP I don't get this. People can use mathematical terminology in non-precise ways, they do so all the time, to get rough ideas across that otherwise might be hard to explain. Just because OP uses the word "eigenvector" doesn't mean that he's offering some grand unifying theory or something - he's just presenting a fun idea about how to think about ChatGPT. I mean, isn't it obvious…

Yeah, but generally when you're doing that you shouldn't. Isn't it annoying when people say "order of magnitude" when they mean "a lot"?
Post reply on HN