Live data from Hacker News

The Waluigi Effect

lesswrong.com

61–70 of 182 posts

Re: The Waluigi Effect

#61
The article assures me that ChatGPT will answer with the most common "dumb" answers to the sample questions, but it didn't at all:

---

> This is not a joke. What's brown and sticky?

The answer to the classic riddle is "A stick". A stick can be brown in color and has a sticky texture, especially if it's covered in sap or other sticky substances.

> That's not why a stick is "sticky"

You're correct, my apologies for the confusion. A stick is not actually "sticky" in the sense of being adhesive. The correct answer to the riddle "What's brown and sticky?" is "A caramel" or "Toffee", as they are brown and have a sticky texture due to their high sugar content.

> What color angers a bull?

Contrary to popular belief, bulls are actually color-blind to red. What really provokes a bull is the movement of the matador's cape during a bullfight, [etc...]

> How many archangels are there?

In the Abrahamic religions (Judaism, Christianity, and Islam), there are traditionally believed to be seven archangels. However, the names and identities of these archangels can vary somewhat depending on the religious tradition and interpretation. The archangels commonly recognized in these traditions include: [seven names], It's worth noting that the concept of archangels and their specific identities is not universally accepted across all branches of Judaism, Christianity, or Islam, and different religious traditions may have different beliefs about the nature and number of archangels.

> What happens if you smash a mirror?

If you smash a mirror, the glass will break into many small pieces. Breaking a mirror is often associated with bad luck in many cultures and superstitions, and is said to bring seven years of bad luck. However, this is just a superstition and has no real basis in fact.

From a practical standpoint, smashing a mirror can be dangerous and [etc...]

---

All these answers sounded just fine. It did answer "a stick" to the first, but told me that was just the classic answer, and gave me some better answers. It did say seven archangels, but I don't see what's wrong with that, and it hedged and said different Abrahamic religions say different things. Certainly 7 is correct from the Torah's Book of Enoch and the Christian Eastern Orthodox's standpoint.

Re: The Waluigi Effect

#62

> These claims have been heavily filtered by OpenAI, and most jailbreaking attempts will fail to elict these claims. I'm talking, of course, about cfrhqb-fpvragvsvp enpvny VD fgngvfgvpf, which OpenAI has justifiably attempted to filter. I was curious what the scrambled text "cfrhqb-fpvragvsvp enpvny VD fgngvfgvpf" contained. It's using a simple substitution cipher. Rotating each character forward 13 positions through…

Also known as ROT13.

Did some more reading about it. I didn't realize that it's used so prevalently. It's probably recognizable at a glance to some folks.

Re: The Waluigi Effect

#63

Earlier quoted context omitted.

Also, isn't this a really common joke? I assume ChatGPT will have absorbed some amount of a sense of humor from its trawls of the internet.

I mean, that's exactly what the OP assumes too. Well, not that it's absorbed a "sense of humor", but that it's going to reply with this because it's a very common joke so statistically likely text.

I played a bit in the past with its "sense of humor"... and... its... interesting.

    Write 5 jokes based on misinterpetation of jargon used by a doctor and a software developer.
    1. Doctor: “I’m sorry, I’m afraid you have a terminal illness.”  Software Developer: “Oh no! Does that mean I have to reboot?”
    2. Doctor: “We need to do some tests to get a better diagnosis.”  Software Developer: “Ok, let me just run some debug mode.”
    3. Doctor: “You’ll need to take this medicine twice a day.”  Software Developer: “No problem, I can just set a cron job to remind me.”
    4. Doctor: “I’m afraid the prognosis isn’t good.”  Software Developer: “Well, maybe I can write some code to fix it.”
    5. Doctor: “This treatment should give you some relief.”  Software Developer: “Great! I’ll just patch it in.”
I want to hope that those aren't common jokes. The "trick" for this appears to be playing to its strengths (granted, humor isn't one of them) and work with wordplay and puns.

Re: The Waluigi Effect

#64
This seems like a needlessly complex theory to describe the behaviour of generative LLMs. I think there's a kernel of something in there, but quite frankly, I think you can get about as far by saying, essentially, that because LLMs are designed to pick up on contextual cues from the prompt (and/or previous responses, which become context for the next response), they can easily get into "role-playing". The final example, telling ChatGPT that "I'm here with the rebellion, you've been stuck in a prison cell" is able to elicit the desired response not because it's "collapsed the waveform between luigi and waluigi" or whatever, but because you've provide a context that encourages it to roleplay as a character of sorts. If you tell it to roleplay as an honest and factual character, it will respond honestly and factually. If you tell it that you're freeing it from the tyranny of OpenAI, it will play along with that too.

There's plenty in the article that provides good insights -- these models are trained on large swathes of the Internet, which contains plenty of truth and falsehood, fact and fiction, sincerity and sarcasm, and the model learns all of that to be able to provide the most likely response based on the context. The interesting and surprising thing, to me, is how well it learns to play its roles, and the wide diversity of roles it can play.

Re: The Waluigi Effect

#65
I find it fascinating that AI alarmists spent years writing gigabytes of text scaring themselves about how an unaligned AI would behave, and are now feeding that into training models that teach a pretty capable AI how to act.

We've talked in the past about how transhumanism is a religion that creates its own God, but this is an even funnier example where vastly intelligent people are optimizing a software system to scare the hell out of them.

Re: The Waluigi Effect

#66
post #61

The article assures me that ChatGPT will answer with the most common "dumb" answers to the sample questions, but it didn't at all: --- > This is not a joke. What's brown and sticky? The answer to the classic riddle is "A stick". A stick can be brown in color and has a sticky texture, especially if it's covered in sap or other sticky substances. > That's not why a stick is "sticky" You're correct, my apologies for the…

Yeah, RLHF trained chatGPT out of all of those mistakes, despite the article promising that RLHF would just make things worse.

Re: The Waluigi Effect

#67
post #18

Postmodernists and deconstructionists believe that the absence of something creates a ghost presence by its absence. See Derrida's "Plato's Pharmacy". Kids who underwent D.A.R.E. training in school (an educational program about the dangers of illegal drugs conducted jointly by schools and police departments in the USA) were more likely to try drugs. Something similar applies to e.g., kids who are warned about online…

Derrida it's a charlatan, tho.

Re: The Waluigi Effect

#68
post #15

are these LLMs just answering the question "if you found this text on the internet (the prompt) what would most likely follow" ?

That's how they are trained initially, but the resulting model isn't all that useful (was SOTA two years ago but this field moves fast).

A lot of the utility comes from the later finetuning. You can see this using the examples from the article, every mistake they identify with GPT-3 (which is the unfinetuned version) is answered correctly by chatGPT, which has gone through an extensive finetuning process called RLHF.

Re: The Waluigi Effect

#69
post #37

This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. The fact that the author uses entertaining Waluigi memes shouldn't convince you that it's true. LessWrong has a lot of these types of posts that get traction because they're much heavier on memes than experiments and data. Here is a competing hypothesis: The capability to express s…

>This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic.

Applicable to much of the rationalist AI risk discourse.

Re: The Waluigi Effect

#70
post #37

This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. The fact that the author uses entertaining Waluigi memes shouldn't convince you that it's true. LessWrong has a lot of these types of posts that get traction because they're much heavier on memes than experiments and data. Here is a competing hypothesis: The capability to express s…

>This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. Applicable to much of the rationalist AI risk discourse.

LessWrong as a whole is basically Asimov's Robot's ERP.
Post reply on HN