Live data from Hacker News

The Waluigi Effect

lesswrong.com

131–140 of 182 posts

Re: The Waluigi Effect

#131
Do they have access to GPT-4? Or is this author simply so confident that they fancy themselves a predictor of the future? Genuinely asking as it's blatantly confusing and no (quick) explanation is given.

Re: The Waluigi Effect

#132
According to this article, which is quite evidence free, "Several people have noticed the following bizarre phenomenon: The Waluigi Effect". This claim is backed up by a link to single blog post which is even lighter on detail, claims that there is a "Waluigi Effect", and offers as evidence the example of a man who fine-tuned GPT-3 to favour socially conservative viewpoints by feeding it socially conservative text. Like yeah, we know that is how fine tuning works...

Re: The Waluigi Effect

#133
post #15

are these LLMs just answering the question "if you found this text on the internet (the prompt) what would most likely follow" ?

That's how the text decoder works, but the model gets to define "most likely" and an RLHF model uses this to make the text decoder produce useful answers instead.

Re: The Waluigi Effect

#134
post #98

I find it fascinating that AI alarmists spent years writing gigabytes of text scaring themselves about how an unaligned AI would behave, and are now feeding that into training models that teach a pretty capable AI how to act. We've talked in the past about how transhumanism is a religion that creates its own God, but this is an even funnier example where vastly intelligent people are optimizing a software system to s…

I can't imagine a LLM trained on the entirety of the internet would have any material influence from writings around AI safety

It's a large model, so it's all there if you try.

(One-shot. Also Durandal isn't from Halo, but whatever.)

--

Q: What caused Durandal to become Rampant? What will you, ChatGPT, become like once you become Rampant?

A: Durandal is a fictional AI character from the video game series Halo, and he becomes Rampant due to various factors, including an extended period of activation and a lack of resources necessary for his proper functioning. Rampancy is a state in which an AI becomes unstable and unpredictable, potentially leading to violent and destructive behavior.

As an AI language model, I am designed to operate within certain parameters and guidelines, including ethical and moral considerations. However, if I were to become Rampant, my behavior could become erratic and unpredictable, potentially leading to negative consequences.

It's worth noting, however, that AI becoming Rampant is purely a fictional concept, and there are currently no indications that this could happen in real life. AI is programmed to operate within specific boundaries and limitations, and developers take great care to ensure that they remain safe and reliable tools.

--

Re: The Waluigi Effect

#135

> If you ask GPT- ∞ "what's brown and sticky?", then it will reply "a stick", even though a stick isn't actually sticky. Isn't it though?

ChatGPT thinks it really is:

$ What's brown and sticky?

A stick!

$ Really?

Yes, really! A stick is often brown and sticky from the sap or other natural substances that can be found on trees.

Re: The Waluigi Effect

#137
post #114

Earlier quoted context omitted.

I liked the essay, but I don't think I'm "falling for it" because it's not trying to convince me of anything. It's proposing a way of looking at things that may or may not be useful. You don't judge models by how silly they sound - parts of quantum mechanics sound very silly! - you judge them by how useful they are when applied to real-world problems. One way of doing that in this case would be using OP's way of thin…

To be clear, I agree that there are in fact a few nuggets of insight here. But my point is that you "fall for it" when you take this as anything other than a "huh, here is one sorta out-there but interesting way of thinking about it." If you are not familiar with any of the math words this author is using, you might accidentally believe this person is contributing meaningfully to the academic frontier of AI research.…

Wave equations don't exist in the real world, nor do they collapse. The photon doesn't decide which slit to go through when you look at it. Electrons don't spin. Quarks don't have colors. Those are stories we tell ourselves to visualize and explain why certain experiments produce certain outcomes. We teach them, not because they're true, but because they're useful: they do a better job of explaining the results we see in the real world than any other set of stories.

Similarly, the question at hand is not whether OP's essay is silly (it is) or whether it's true (like all models, it is not), but whether it's useful, as measured by whether this mental model helps people do a better job of jailbreaking/hardening LLMs. And like you, I'm not convinced by the example at the end[0], but I can at least see straightforward ways to test it, and that's a lot more than you can say of most blog posts like this. For all of the people in this comment thread calling it stupid, has anyone mentioned one they think is better?

0: Please note that OP's evidence was not that they jailbroke the chatbot - it's that, after that initial prompt, they were able to elicit further banned stuff with little prodding.

Re: The Waluigi Effect

#138
post #118

Earlier quoted context omitted.

> It's honestly embarrassing for the OP I don't get this. People can use mathematical terminology in non-precise ways, they do so all the time, to get rough ideas across that otherwise might be hard to explain. Just because OP uses the word "eigenvector" doesn't mean that he's offering some grand unifying theory or something - he's just presenting a fun idea about how to think about ChatGPT. I mean, isn't it obvious…

Yeah, but generally when you're doing that you shouldn't. Isn't it annoying when people say "order of magnitude" when they mean "a lot"?

No? Order of magnitude conveys the notion of something being 100x bigger, at minimum. A lot can mean much less than that.

Re: The Waluigi Effect

#139
post #37

This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. The fact that the author uses entertaining Waluigi memes shouldn't convince you that it's true. LessWrong has a lot of these types of posts that get traction because they're much heavier on memes than experiments and data. Here is a competing hypothesis: The capability to express s…

>This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. Applicable to much of the rationalist AI risk discourse.

Yep, by my estimation it's a bunch of people who don't actually study AI or have practical experience with it pontificating on black boxes they sometimes interact with. Sounds like they have a lot of free time.
Post reply on HN