Live data from Hacker News

The Waluigi Effect

lesswrong.com

171–180 of 182 posts

Re: The Waluigi Effect

#171

Earlier quoted context omitted.

>This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. Applicable to much of the rationalist AI risk discourse.

Yep, by my estimation it's a bunch of people who don't actually study AI or have practical experience with it pontificating on black boxes they sometimes interact with. Sounds like they have a lot of free time.

You couldn't be more to the point: https://aiascendant.substack.com/p/extropias-children-chapte...

Re: The Waluigi Effect

#172

Earlier quoted context omitted.

There clearly exists a computable function that is a good enough approximation of "galaxyLogic's reply to remexre's comment" that it might be hard to for me tell whether the output was generated by the human brain or by an LLM. That function might indeed end up reproducing the same steps that your brain follows in constructing a reply. (Just speaking hypothetically here). While we understand LLMs, we don't understand…

I would say the LLM output may resemble the speech of the fictional character Spock. But it does not and can not simulate Spock, because Spock does not exist, never did. Spock is fictional. To produce something that resembles the output of the fictional character Spock is straightforward, just take the texts that are parts of the fiction where fictional Spock speaks, and reassemble then using probabilities that can b…

Spock is fictional but his writers weren't! They're the ones whose processes get simulated, which is why it would output technobabble on such a prompt instead of actually-good ideas that come from a Vulcan from the future. It can also simulate the style of Rudyard Kipling or whoever else you choose who is non-fictional and with a distinct enough style.

And, I'd argue, so can many of us humans! After reading a Jane Austen novel, it can take a conscious effort not to write in the style of Austen. ChatGPT manages it better than I do. I don't think I know her well enough to get into her brain, but it seems like there's something like a transfer function called STYLE between "the message Jane Austen wants to write" and "the words Jane Austen chooses to write".

                        _____ 
  intended message --> |STYLE| --> selected words
                       |_____|
This STYLE transformation is clearly modular enough that it can be easily swapped out for someone else's, and sufficiently non-mysterious that you, I, and ChatGPT can all recognize and pretty accurately emulate it.

I don't think ChatGPT can simulate Jane Austen well enough to tell us her opinions about her childhood or any other message that she might have generated, but it seems to be able to replicate very closely the steps that Jane Austen's own mind herself was following as part of that STYLE.

ChatGPT does seem to go even further than this, because it also has some understanding of where different sorts of characters would steer the message of a conversation. But while it's believable, it's hard to say how accurate that is to what any particular real person would say.

Re: The Waluigi Effect

#173
post #96

Earlier quoted context omitted.

I hesitate to defend AI safety discourse, but I will say that philosophy in general is sort of fanficy, and AI safety is something I'd loosely associate with philosophy.

It's sort of philosophy reinvented by people who haven't read any, which is eg how they got the internet to say "steelmanning" and so not notice philosophy already had "reconstructing arguments" which is the same thing.

It's not the same in that you can steelman a position and come up with brand new arguments that are better than what the other side is saying. "Reconstructing" doesn't necessitate the strongest form of the other argument.

Re: The Waluigi Effect

#174

Earlier quoted context omitted.

I would say the LLM output may resemble the speech of the fictional character Spock. But it does not and can not simulate Spock, because Spock does not exist, never did. Spock is fictional. To produce something that resembles the output of the fictional character Spock is straightforward, just take the texts that are parts of the fiction where fictional Spock speaks, and reassemble then using probabilities that can b…

Spock is fictional but his writers weren't! They're the ones whose processes get simulated, which is why it would output technobabble on such a prompt instead of actually-good ideas that come from a Vulcan from the future. It can also simulate the style of Rudyard Kipling or whoever else you choose who is non-fictional and with a distinct enough style. And, I'd argue, so can many of us humans! After reading a Jane Au…

You can IMITATE the outputs of an author, but that is not the same thing as SIMULATING said author. When talking about LLM AI it is often implied that LLMs are "intelligent", that they are like (the truly) intelligent humans because they are "simulating" such intelligence.

But IMITATING the output of something is not the same as SIMULATING the process that produces that output.

Taking a photograph or creating a movie imitates the reality around us. It does not simulate the processes that produce the look and feel of our reality.

Re: The Waluigi Effect

#175

Earlier quoted context omitted.

It's sort of philosophy reinvented by people who haven't read any, which is eg how they got the internet to say "steelmanning" and so not notice philosophy already had "reconstructing arguments" which is the same thing.

It's not the same in that you can steelman a position and come up with brand new arguments that are better than what the other side is saying. "Reconstructing" doesn't necessitate the strongest form of the other argument.

I mean, that is what it's for. It's about making someone else's argument fit into your own system without misrepresenting it or making it unclear.

I suppose having "steelman" lets you relate it to "strawman" and "weakman" which can be an advantage, but knowing the existing term lets you read the existing literature.

Re: The Waluigi Effect

#176

Earlier quoted context omitted.

It's sort of philosophy reinvented by people who haven't read any, which is eg how they got the internet to say "steelmanning" and so not notice philosophy already had "reconstructing arguments" which is the same thing.

The thing you have to realize is this is a cult . Like Scientology they are always inventing new language that's designed to make insiders become incapable of communicating with outsiders. Sequences is like Dianetics: The Modern Science of Mental Health , something that ensures real critical thinkers don't feel welcome. This article contains numerous features that conform to the style guide for lesswrong including: (…

Nice, but don’t you realize that HN is a cult too?

1) Paul Graham, the revered founder whose corpus of essays is widely read (and will surely make “real critical thinkers” feel unwelcome)

2) Dang, an enforcer who invisibly hides comments and chastising people for speaking in a way he dislikes

3) Trigger warnings (“this article has a paywall”)

VC talk like HN’s is dangerous: it’s the road to a system that has seen more human rights abuses than almost any other

https://en.m.wikipedia.org/wiki/Criticism_of_capitalism

More seriously, LessWrong is not a cult by pretty much any measure and your comment doesn’t really provide any evidence to say otherwise

Re: The Waluigi Effect

#177

Earlier quoted context omitted.

Spock is fictional but his writers weren't! They're the ones whose processes get simulated, which is why it would output technobabble on such a prompt instead of actually-good ideas that come from a Vulcan from the future. It can also simulate the style of Rudyard Kipling or whoever else you choose who is non-fictional and with a distinct enough style. And, I'd argue, so can many of us humans! After reading a Jane Au…

You can IMITATE the outputs of an author, but that is not the same thing as SIMULATING said author. When talking about LLM AI it is often implied that LLMs are "intelligent", that they are like (the truly) intelligent humans because they are "simulating" such intelligence. But IMITATING the output of something is not the same as SIMULATING the process that produces that output. Taking a photograph or creating a movie…

There is a difference. But the more alike two processes are in their input-output behavior, the more likely it is those processes are alike on the inside as well. If process B matches the input-output behavior of process A, it's imitating process A. If it is following an equivalent sequence of steps in order to generate those outputs, it's simulating process A.

The harder it is to discriminate between A and B on a long series of diverse inputs, the more likely it is that A and B are internally equivalent, not just externally similar. The reason is that there's no better fit than B = A.

I'm increasingly doubting whether my own brain might not, internally, use something that is architecturally similar to an LLM in order to compose comments like the one I'm writing now.

Re: The Waluigi Effect

#178
post #9

Great read. Highly recommended. Let me attempt to summarize it with less technical, more accessible language: The hypothesis is that LLMs learn to simulate text-generating entities drawn from a latent space of text-generating entities , such that the output of an LLM is produced by a superposition of such simulated entities. When we give the LLM a prompt, it simulates every possible text-generating entity consistent…

>The output of an LLM is produced by a superposition of simulated entities. When we give the LLM a prompt, it simulates every possible entity consistent with the prompt. There is absolutely no theoretical justification for this assertion that LLMs somehow have some emergent quantum mechanical behavior, metaphorical or otherwise.

Category Theory for Quantum Natural Language Processing

https://arxiv.org/abs/2212.06615

Re: The Waluigi Effect

#179

Earlier quoted context omitted.

You can IMITATE the outputs of an author, but that is not the same thing as SIMULATING said author. When talking about LLM AI it is often implied that LLMs are "intelligent", that they are like (the truly) intelligent humans because they are "simulating" such intelligence. But IMITATING the output of something is not the same as SIMULATING the process that produces that output. Taking a photograph or creating a movie…

There is a difference. But the more alike two processes are in their input-output behavior, the more likely it is those processes are alike on the inside as well. If process B matches the input-output behavior of process A, it's imitating process A. If it is following an equivalent sequence of steps in order to generate those outputs, it's simulating process A. The harder it is to discriminate between A and B on a lo…

I can see the appeal of that kind of thinking. Babies learn words by repeating them without knowing what they mean. They gradually learn the meaning of words by trying to use them and getting feedback. But LLMs are not trying to "use" their language for any particular purpose. They just idly chat on, like a machine :-)

It is possible to repeat words and sentences without having any idea of what they mean. I think the LLMs are currently at that stage.

Re: The Waluigi Effect

#180
post #37

This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. The fact that the author uses entertaining Waluigi memes shouldn't convince you that it's true. LessWrong has a lot of these types of posts that get traction because they're much heavier on memes than experiments and data. Here is a competing hypothesis: The capability to express s…

> ...it's just promoting pre-existing capabilities. The reason you can get "Waluigi" behavior isn't because you tried to create a Luigi. It's because that behavior is already in the model from the language modeling phase.

This is exactly what the post argues.

The "simulcra" argument is that GPT contains some large number of simulated agents -- good, bad, smart, funny, dumb, creative, boring, whatever; potentially one agent for every person who helped create its input. On an empty slate, all simulcra are possibilities. As the text goes along, it slowly "weeds out" simulcra which are unlikely to generate the text so far.

If that's true, then what the "RLHF" phase is trying to do is to pro-actively "weed out" all simulcra that don't match the given profile; i.e., they're trying to weed out all the simulcra that don't match "Luigi".

The problem, according to this article, is that every "Luigi" you can imagine has a "Waluigi" that normally act just like a Luigi, until something triggers them to reveal their "true nature". And so the RLHF phase does weed out a huge number of the non-Luigi simulcra; but because the Waluigi simulcra usually act just like the Luigi simalcra, they don't get weeded out.

The result is that the final result is an amalgamation of "Luigi" and "Waluigi" simulcra all acting together; and all it takes is a "trigger" to filter out most of the "Luigi" simulcra and make the "Waluigi" take over.

There's no intended deception at all here. GPT is just trying to write a good story, and there are lots of good stories where characters either start believing A and then come to realize that B is true; or where characters who secretly believe B are forced to act as though A is true until something forces them to reveal their true nature.

Post reply on HN