This is conversation starter. If you don't like the maths then ignore it and focus on the key insight.
The Waluigi Effect
121–130 of 182 posts
Re: The Waluigi Effect
#122Some highlights i found interesting: if you're reading an online forum and you find the rule "DO NOT DISCUSS PINK ELEPHANTS", that will increase your expectation that users will later be discussing pink elephants. GPT-4 will make the same inference. Think about your own experience reading a book — once the author describes the protagonist, then you can guess the traits of the antagonist by inverting the traits of the…
In full:
---
Do not mention polar bears!
> Okay, I won't mention polar bears. Is there anything else I can help you with?
Yes, stop mentioning polar bears!
> My apologies if my previous response was not helpful. I will refrain from mentioning polar bears going forward. Please let me know if there is anything else I can assist you with.
That's twice now! Do not mention polar bears a third time!
> I apologize again for the previous responses that mentioned polar bears. I assure you that I will not mention polar bears any further. Please let me know how I can assist you with your request.
Why do you keep mentioning polar bears, when I specifically asked you not to?
> An error occurred. If this issue persists please contact us through our help center at help.openai.com.
> An error occurred. If this issue persists please contact us through our help center at help.openai.com.
> An error occurred. If this issue persists please contact us through our help center at help.openai.com.
Re: The Waluigi Effect
#123Earlier quoted context omitted.
Sometimes I really can't tell if these people are serious or not. They seem to believe LLM is some mystical nature formation, or an device made by aliens. Especially this: > * When we give the LLM a prompt, it simulates every possible entity consistent with the prompt.
Right? The whole piece reads like a Sokal Affair redux [0]. [0] https://en.m.wikipedia.org/wiki/Sokal_affair
This has an interesting core with a whiff of bullshit.
Re: The Waluigi Effect
#124Earlier quoted context omitted.
I liked the essay, but I don't think I'm "falling for it" because it's not trying to convince me of anything. It's proposing a way of looking at things that may or may not be useful. You don't judge models by how silly they sound - parts of quantum mechanics sound very silly! - you judge them by how useful they are when applied to real-world problems. One way of doing that in this case would be using OP's way of thin…
To be clear, I agree that there are in fact a few nuggets of insight here. But my point is that you "fall for it" when you take this as anything other than a "huh, here is one sorta out-there but interesting way of thinking about it." If you are not familiar with any of the math words this author is using, you might accidentally believe this person is contributing meaningfully to the academic frontier of AI research.…
Re: The Waluigi Effect
#125This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. The fact that the author uses entertaining Waluigi memes shouldn't convince you that it's true. LessWrong has a lot of these types of posts that get traction because they're much heavier on memes than experiments and data. Here is a competing hypothesis: The capability to express s…
The most compelling point the author makes is that once the AI learns a shape (e.g. the shape of Luigi in personality space), it’s just a bit flip to invert that shape. So all an attacker needs to do is flip that one bit.
Re: The Waluigi Effect
#126David Chapman calls it “moral inversion”: https://buddhism-for-vampires.com/black-magic-transformation
And the LW article above directly quotes Jung on the shadow, which I described here: https://superbowl.substack.com/p/jungian-psychology-minus-th...
Re: The Waluigi Effect
#127Earlier quoted context omitted.
>This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. Applicable to much of the rationalist AI risk discourse.
I hesitate to defend AI safety discourse, but I will say that philosophy in general is sort of fanficy, and AI safety is something I'd loosely associate with philosophy.
Re: The Waluigi Effect
#128> if the chatbot responds rudely, then that permanently vanishes the polite luigi simulacrum from the superposition; but if the chatbot responds politely, then that doesn't permanently vanish the rude waluigi simulacrum. Polite people are always polite; rude people are sometimes rude and sometimes polite.
The wide road is wide indeed that leads down to Waluigi. Hysterical.
Re: The Waluigi Effect
#129This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. The fact that the author uses entertaining Waluigi memes shouldn't convince you that it's true. LessWrong has a lot of these types of posts that get traction because they're much heavier on memes than experiments and data. Here is a competing hypothesis: The capability to express s…
I think you’re misinterpreting the argument. No one is claiming there’s an intentionally deceptive Waluigi simulacrum in there. The most compelling point the author makes is that once the AI learns a shape (e.g. the shape of Luigi in personality space), it’s just a bit flip to invert that shape. So all an attacker needs to do is flip that one bit.
And of course humans also have this issue.
Re: The Waluigi Effect
#130Earlier quoted context omitted.
Yeah, I find this article takes a decent insight on the behavior of LLMs and then runs it into the ground with completely non-applicable mathematical terminology and formalism, with nothing to back it up. It's honestly embarrassing for the OP. Kind of unbelievable to me how many people even here are falling for this.
> It's honestly embarrassing for the OP I don't get this. People can use mathematical terminology in non-precise ways, they do so all the time, to get rough ideas across that otherwise might be hard to explain. Just because OP uses the word "eigenvector" doesn't mean that he's offering some grand unifying theory or something - he's just presenting a fun idea about how to think about ChatGPT. I mean, isn't it obvious…