The Waluigi Effect
11–20 of 182 posts
Re: The Waluigi Effect
#12It's fascinating to see that, as they develop on massive corpuses of human output, neural networks are rapidly moving from something which can be analyzed in terms of math and computer science, from something which needs to be analyzed using the "softer" sciences of psychology. It's something I think people are not ready for (notice the comments in here already griping that this is unverifiable speculation - which is true, in a sense, but we don't really have any other choice).
Re: The Waluigi Effect
#13Great read. Highly recommended. Let me attempt to summarize it with less technical, more accessible language: The hypothesis is that LLMs learn to simulate text-generating entities drawn from a latent space of text-generating entities , such that the output of an LLM is produced by a superposition of such simulated entities. When we give the LLM a prompt, it simulates every possible text-generating entity consistent…
There is absolutely no theoretical justification for this assertion that LLMs somehow have some emergent quantum mechanical behavior, metaphorical or otherwise.
Re: The Waluigi Effect
#14Great read. Highly recommended. Let me attempt to summarize it with less technical, more accessible language: The hypothesis is that LLMs learn to simulate text-generating entities drawn from a latent space of text-generating entities , such that the output of an LLM is produced by a superposition of such simulated entities. When we give the LLM a prompt, it simulates every possible text-generating entity consistent…
Is he really evil though? I thought all he did was play tennis and golf and drive a go-kart.
Re: The Waluigi Effect
#15Re: The Waluigi Effect
#16Great read. Highly recommended. Let me attempt to summarize it with less technical, more accessible language: The hypothesis is that LLMs learn to simulate text-generating entities drawn from a latent space of text-generating entities , such that the output of an LLM is produced by a superposition of such simulated entities. When we give the LLM a prompt, it simulates every possible text-generating entity consistent…
>The output of an LLM is produced by a superposition of simulated entities. When we give the LLM a prompt, it simulates every possible entity consistent with the prompt. There is absolutely no theoretical justification for this assertion that LLMs somehow have some emergent quantum mechanical behavior, metaphorical or otherwise.
Also, I'm simplifying things a lot to make them accessible.
The OP goes into a lot more detail.
I highly recommend you read it.
Re: The Waluigi Effect
#17This is what the Waluigi effect is, since it isn't described at the top: > The Waluigi Effect: After you train an LLM to satisfy a desirable property P, then it's easier to elicit the chatbot into satisfying the exact opposite of property P. Basically, the chatbot will often do the opposite of what you say.
IMO it's more like if you tell the LLM to never talk about Pink Elephants, it will become easier to get it to talk about Pink Elephants later. (It is easier to get an anti-Pink Elephant model to talk about Pink Elephants than it is to get a neutral model to talk about Pink Elephants)
I have similar thoughts about swearing being kept alive by teaching children not to say "bad" words and various kinds of bigotry being amplified at this point by people trying to fight against it.
Re: The Waluigi Effect
#18Kids who underwent D.A.R.E. training in school (an educational program about the dangers of illegal drugs conducted jointly by schools and police departments in the USA) were more likely to try drugs. Something similar applies to e.g., kids who are warned about online porn: the warning stokes their curiosity.
"If you have a pink duck and a pink lion and a green duck, ask yourself where the green lion has gotten to." --Alan G. Carter
Re: The Waluigi Effect
#19are these LLMs just answering the question "if you found this text on the internet (the prompt) what would most likely follow" ?
How you train a model and the inference jobs it can do don't necessarily have to be the same.
Re: The Waluigi Effect
#20Great read. Highly recommended. Let me attempt to summarize it with less technical, more accessible language: The hypothesis is that LLMs learn to simulate text-generating entities drawn from a latent space of text-generating entities , such that the output of an LLM is produced by a superposition of such simulated entities. When we give the LLM a prompt, it simulates every possible text-generating entity consistent…
> * When we give the LLM a prompt, it simulates every possible entity consistent with the prompt.