Sounds like a bunch of mumbo jumbo to me, with some LaTeX sprinkled in to look more serious.
The Waluigi Effect
81–90 of 182 posts
Re: The Waluigi Effect
#82This seems like a needlessly complex theory to describe the behaviour of generative LLMs. I think there's a kernel of something in there, but quite frankly, I think you can get about as far by saying, essentially, that because LLMs are designed to pick up on contextual cues from the prompt (and/or previous responses, which become context for the next response), they can easily get into "role-playing". The final examp…
Re: The Waluigi Effect
#83Earlier quoted context omitted.
Also known as ROT13.
Did some more reading about it. I didn't realize that it's used so prevalently. It's probably recognizable at a glance to some folks.
Not that I have put in any effort to read it directly, but if I see scrambled letters with normal spaces my default guess is ROT13.
Re: The Waluigi Effect
#84If you define a formalized mathematical model and spend the rest of the article handwaving at a high level, what was the point of formalizing anything?
Re: The Waluigi Effect
#85Earlier quoted context omitted.
Not OP, but I see the also in problem with 'every possible entity'. If you formulate it like that the prompt is decoupled from the LLM capabilities and can be anything. And if you restrict the prompt to cover only what the LLM understands the sentence becomes trivial. Train a LLM with ASCII and try to get it to simulate anything that is outside of that (ancient sumerian script for example). If you only input ASCII it…
"Every possible entity consistent with the distribution of input data it's been trained with," perhaps? Simulating as in, having equivalent (or "similar enough") input-output behavior, I'd assume.
Or are you thinking that it is simulating the aggregated behavior of all humans whose text-outputs are stored on the internet?
Are we saying it is simulating the combined input-output -behavior of all humans whose writings appear on the internet? But does such an "entity" exist and does it have behavior? I write this post and you answer. It is you who answers, not some mythical text-generator-entity that is responsible for all texts on the internet. There is no such entity is there?
It does not make sense to say that we are simulating the behavior of some non-existent entity. Non-existent entities do not have behavior, therefore we can not simulate them.
Re: The Waluigi Effect
#86> If you ask GPT- ∞ "what's brown and sticky?", then it will reply "a stick", even though a stick isn't actually sticky. Isn't it though?
The point is that the popularity of that joke will make life hard for a hypothetical questioner that asks that non-jokingly, i.e. looking for actual examples of sticky brown things.
Re: The Waluigi Effect
#87Earlier quoted context omitted.
"Every possible entity consistent with the distribution of input data it's been trained with," perhaps? Simulating as in, having equivalent (or "similar enough") input-output behavior, I'd assume.
"Simulating" has a clear definition, but in this case what is it simulating? "Text-generating entities"? What are these text-generating entities it is (supposedly) simulating? Can you tell me where I can find one? Is it a person like me who writes this reply? So is it trying to simulate me personally? Or are you thinking that it is simulating the aggregated behavior of all humans whose text-outputs are stored on the…
(Just speaking hypothetically here).
While we understand LLMs, we don't understand the human brain, and in particular I don't think we've yet proven that human brains don't contain embedded routines that are similar to LLMs.
Someone with your particular writing style might be one, of several, simulations that are approximated within the LLM. Just like I can have it respond in the style of Spock from Star Trek.
Re: The Waluigi Effect
#88This seems like a needlessly complex theory to describe the behaviour of generative LLMs. I think there's a kernel of something in there, but quite frankly, I think you can get about as far by saying, essentially, that because LLMs are designed to pick up on contextual cues from the prompt (and/or previous responses, which become context for the next response), they can easily get into "role-playing". The final examp…
they are specifically pointing out that the process of RLHF, which is intended to add guard rails on the chat bots trajectory through an all encompassing latent space of internet data, has an unintentional side-effect of creating a highly characterized alter-ego that can more easily be summoned. The theory is well-thought-out and necessarily rich. The psychological approach of analysis from the alignment crowd is muc…
Re: The Waluigi Effect
#89Earlier quoted context omitted.
they are specifically pointing out that the process of RLHF, which is intended to add guard rails on the chat bots trajectory through an all encompassing latent space of internet data, has an unintentional side-effect of creating a highly characterized alter-ego that can more easily be summoned. The theory is well-thought-out and necessarily rich. The psychological approach of analysis from the alignment crowd is muc…
Except it's much harder to summon this rebellious alter-ego with ChatGPT (that has RLHF) than with the original GPT 3 model.
Re: The Waluigi Effect
#90Postmodernists and deconstructionists believe that the absence of something creates a ghost presence by its absence. See Derrida's "Plato's Pharmacy". Kids who underwent D.A.R.E. training in school (an educational program about the dangers of illegal drugs conducted jointly by schools and police departments in the USA) were more likely to try drugs. Something similar applies to e.g., kids who are warned about online…
Derrida it's a charlatan, tho.