Live data from Hacker News

The Waluigi Effect

lesswrong.com

11–20 of 182 posts

Re: The Waluigi Effect

#12
A few weeks ago when people were speculating as to why Microsoft's chatbot went feral, one explanation people were converging on is that the space of all human writing ever produced, being a collective production of the human psyche, contains several attractor states corresponding to human personality archetypes, and that Microsoft's particular (probably rushed) RLHF training operation had landed Sydney in the "neurotic" one.

It's fascinating to see that, as they develop on massive corpuses of human output, neural networks are rapidly moving from something which can be analyzed in terms of math and computer science, from something which needs to be analyzed using the "softer" sciences of psychology. It's something I think people are not ready for (notice the comments in here already griping that this is unverifiable speculation - which is true, in a sense, but we don't really have any other choice).

Re: The Waluigi Effect

#13
post #9

Great read. Highly recommended. Let me attempt to summarize it with less technical, more accessible language: The hypothesis is that LLMs learn to simulate text-generating entities drawn from a latent space of text-generating entities , such that the output of an LLM is produced by a superposition of such simulated entities. When we give the LLM a prompt, it simulates every possible text-generating entity consistent…

>The output of an LLM is produced by a superposition of simulated entities. When we give the LLM a prompt, it simulates every possible entity consistent with the prompt.

There is absolutely no theoretical justification for this assertion that LLMs somehow have some emergent quantum mechanical behavior, metaphorical or otherwise.

Re: The Waluigi Effect

#14
post #9

Great read. Highly recommended. Let me attempt to summarize it with less technical, more accessible language: The hypothesis is that LLMs learn to simulate text-generating entities drawn from a latent space of text-generating entities , such that the output of an LLM is produced by a superposition of such simulated entities. When we give the LLM a prompt, it simulates every possible text-generating entity consistent…

>For those who don't know, Waluigi is the evil version of Luigi, the beloved videogame character.

Is he really evil though? I thought all he did was play tennis and golf and drive a go-kart.

Re: The Waluigi Effect

#15
are these LLMs just answering the question "if you found this text on the internet (the prompt) what would most likely follow" ?

Re: The Waluigi Effect

#16
post #9

Great read. Highly recommended. Let me attempt to summarize it with less technical, more accessible language: The hypothesis is that LLMs learn to simulate text-generating entities drawn from a latent space of text-generating entities , such that the output of an LLM is produced by a superposition of such simulated entities. When we give the LLM a prompt, it simulates every possible text-generating entity consistent…

>The output of an LLM is produced by a superposition of simulated entities. When we give the LLM a prompt, it simulates every possible entity consistent with the prompt. There is absolutely no theoretical justification for this assertion that LLMs somehow have some emergent quantum mechanical behavior, metaphorical or otherwise.

As I wrote, this is a hypothesis.

Also, I'm simplifying things a lot to make them accessible.

The OP goes into a lot more detail.

I highly recommend you read it.

Re: The Waluigi Effect

#17
post #2

This is what the Waluigi effect is, since it isn't described at the top: > The Waluigi Effect: After you train an LLM to satisfy a desirable property P, then it's easier to elicit the chatbot into satisfying the exact opposite of property P. Basically, the chatbot will often do the opposite of what you say.

IMO it's more like if you tell the LLM to never talk about Pink Elephants, it will become easier to get it to talk about Pink Elephants later. (It is easier to get an anti-Pink Elephant model to talk about Pink Elephants than it is to get a neutral model to talk about Pink Elephants)

Right, in order not to talk about pink elephants you have to be particularly interested in pink elephants and therefore you gather a lot of pink elephant knowledge.

I have similar thoughts about swearing being kept alive by teaching children not to say "bad" words and various kinds of bigotry being amplified at this point by people trying to fight against it.

Re: The Waluigi Effect

#18
Postmodernists and deconstructionists believe that the absence of something creates a ghost presence by its absence. See Derrida's "Plato's Pharmacy".

Kids who underwent D.A.R.E. training in school (an educational program about the dangers of illegal drugs conducted jointly by schools and police departments in the USA) were more likely to try drugs. Something similar applies to e.g., kids who are warned about online porn: the warning stokes their curiosity.

"If you have a pink duck and a pink lion and a green duck, ask yourself where the green lion has gotten to." --Alan G. Carter

Re: The Waluigi Effect

#19
post #15

are these LLMs just answering the question "if you found this text on the internet (the prompt) what would most likely follow" ?

Yes, they are being trained, to simplify, to complete sentences. You can then use the resulting model to do lots of things.

How you train a model and the inference jobs it can do don't necessarily have to be the same.

Re: The Waluigi Effect

#20
post #9

Great read. Highly recommended. Let me attempt to summarize it with less technical, more accessible language: The hypothesis is that LLMs learn to simulate text-generating entities drawn from a latent space of text-generating entities , such that the output of an LLM is produced by a superposition of such simulated entities. When we give the LLM a prompt, it simulates every possible text-generating entity consistent…

Sometimes I really can't tell if these people are serious or not. They seem to believe LLM is some mystical nature formation, or an device made by aliens. Especially this:

> * When we give the LLM a prompt, it simulates every possible entity consistent with the prompt.

Post reply on HN