Live data from Hacker News

The Waluigi Effect

lesswrong.com

151–160 of 182 posts

Re: The Waluigi Effect

#151

Earlier quoted context omitted.

Yeah, but generally when you're doing that you shouldn't. Isn't it annoying when people say "order of magnitude" when they mean "a lot"?

No? Order of magnitude conveys the notion of something being 100x bigger, at minimum. A lot can mean much less than that.

You mean 10x bigger. One order of magnitude is 10x.

Re: The Waluigi Effect

#152

I would question the assumption that there is a simulacrum of anything in a LLM, not even implicit. Any simulacrum, identity, self-consistency etc. is a projection of the "reader", i.e. user. (I guess it is an interesting philosophical question whether a convincing presentation of a simulation of a mind is a mind, or at least an acceptable simulation. One meta level higher as the turing test, so to speak. If so, I'm…

Do you think an ant have a subjective experience? If not, why? If so, why wouldn't a computer, or parts of a computer?

Based on that reasoning, why wouldn't an economy or a corporation have subjective experience?

Re: The Waluigi Effect

#153
post #16

Earlier quoted context omitted.

As I wrote, this is a hypothesis . Also, I'm simplifying things a lot to make them accessible. The OP goes into a lot more detail. I highly recommend you read it.

I did read it. The whole article reads like someone trying to make a loose conjecture appear quantitatively rigorous by abusing terminology from physics, statistics, and chaos theory (among other quantitative fields). For example, >the superposition is unlikely to collapse to the luigi simulacrum because there is no behaviour which is likely for luigi but very unlikely for waluigi. Recall that the waluigi is pretendi…

The K-L divergence is relevant there, even though I'm pretty sure that that "formally" comment is meant as a joke and not serious.

(wikipedia) "A simple interpretation of the KL divergence of P from Q is the expected excess surprise from using Q as a model when the actual distribution is P.

The sentence you quoted posits that whenever the LMM is in a state where it is "simulating" a nice and helpful person, the simulation is also consistent with an insane, violent person that's currently pretending to be nice, but not the other way around.

The author isn't talking about the loss or error of predicting individual tokens. If you look the larger scale behaviour, predicting, e.g., a scalar niceness value of the response based on one of two "modes" that you assume the LLM is currently in (either Waluigi or Luigi), then you'll be less surprised if Waluigi acts like Luigi than the other way around.

The probability distribution of niceness when assuming the LLM is in a state of "Luigi" would have a high mean and low variance, while the distribution for Waluigi would have a lower mean but a higher variance.

Thus, the KL divergence of Waluigi (that is, the probability distribution of niceness you'd predict when assuming the model is in Waluigi mode) from Luigi would be high, while the other way around `KL(Luigi, Waluigi)` would be low.

It should be easy to construct an example with concrete values using two normal probablility distributions.

Re: The Waluigi Effect

#154

Earlier quoted context omitted.

That assumption does seem pretty unlikely a priori. After all, the OpenAI folks added RLHF to GPT-3, presumably did some testing, and then opened it to the public. If the testing noticed more antisocial behavior after adding RLHF, presumably that would not have been the version they opened up. One might argue that the model was able to successfully hide the antisocial behavior from the testers, but that seems unlikel…

Why do you think it's unlikely? Internal testing with a few alpha testers and some automated testing is useful, but lots of bugs are only found in wider testing or in production. Chatbot conversations are open-ended, so it's not surprising to me that when you get tens or hundreds of thousands of people doing testing then they're going to find more weird behaviors, particularly since they're actively trying to "break"…

I mean, sure, it's going to expose more weird behaviors with a wider audience looking at it. The core problem is that it's so easy to get ChatGPT to start exhibiting weird behaviors that it would be surprising if the testers just never ran into them. Remember, the internal testing is actively trying to break things too, and they can use knowledge of the internals and past versions to do so.

Also, the assumption I find dubious is that RLHF results in more antisocial behavior than not using it. Both versions would have been tested, so OpenAI would've had a baseline from testing the prior version with equal or fewer resources. Equal or greater rigor, and you'd expect them to open it up only if they found fewer flaws.

Re: The Waluigi Effect

#155

I would question the assumption that there is a simulacrum of anything in a LLM, not even implicit. Any simulacrum, identity, self-consistency etc. is a projection of the "reader", i.e. user. (I guess it is an interesting philosophical question whether a convincing presentation of a simulation of a mind is a mind, or at least an acceptable simulation. One meta level higher as the turing test, so to speak. If so, I'm…

> Especially there is no world-model, and no inner state.

Some people argue otherwise[1]. It’s an interesting debate.

[1] https://twitter.com/random_walker/status/1631502179323215872...

Re: The Waluigi Effect

#157
I've read the article and frankly I don't have the expertise on LLMs to discuss whether it lands on something worth investigating further. What I do notice in the comments here though, are:

1) People nitpicking about the use of mathematical ideas in a loose manner as if every person trying to understand some phenomenon must only open their mouth if they have a watertight theory or shut their mouth otherwise.

2) Getting hung up on the use of the luigi metaphors rather than using it as the basis for a constructive criticism that actually adds to the conversation in an interesting manner.

3) A general snarky attitude towards people exploring ideas on their own. I get it, you might have some expertise that others lack but you're forgetting that you've already made the thousands of mistakes to get to where you are. Do others the courtesy of not judging when they attempt the same.

Re: The Waluigi Effect

#158
post #95

While the formality is way overwrought (and ChatGPT is not creating any "simulacra" of characters), I think the overall point is correct that language models are trained on stories and other human writing, and inversion is a very common plot point in stories and human writing in general, if only because contradicting expectations is more interesting (e.g. "man bites dog"). We also less commonly see exposition that is…

The simulacra theory is an apt one, see https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators

Particularly, as noted by David Chalmers:

> What pops out of self-supervised predictive training is noticeably not a classical agent. Shortly after GPT-3’s release, David Chalmers lucidly observed that the policy’s relation to agents is like that of a “chameleon” or “engine”:

>> GPT-3 does not look much like an agent. It does not seem to have goals or preferences beyond completing text, for example. It is more like a chameleon that can take the shape of many different agents. Or perhaps it is an engine that can be used under the hood to drive many agents. But it is then perhaps these systems that we should assess for agency, consciousness, and so on.[6]

Re: The Waluigi Effect

#159

Earlier quoted context omitted.

"Simulating" has a clear definition, but in this case what is it simulating? "Text-generating entities"? What are these text-generating entities it is (supposedly) simulating? Can you tell me where I can find one? Is it a person like me who writes this reply? So is it trying to simulate me personally? Or are you thinking that it is simulating the aggregated behavior of all humans whose text-outputs are stored on the…

There clearly exists a computable function that is a good enough approximation of "galaxyLogic's reply to remexre's comment" that it might be hard to for me tell whether the output was generated by the human brain or by an LLM. That function might indeed end up reproducing the same steps that your brain follows in constructing a reply. (Just speaking hypothetically here). While we understand LLMs, we don't understand…

I would say the LLM output may resemble the speech of the fictional character Spock. But it does not and can not simulate Spock, because Spock does not exist, never did. Spock is fictional.

To produce something that resembles the output of the fictional character Spock is straightforward, just take the texts that are parts of the fiction where fictional Spock speaks, and reassemble then using probabilities that can be calculated by statistically analyzing those texts. That is what LLMs are doing, right? And results can be quite surprising. I assume people were similarly impressed when they first saw movies.

But LLMs are not simulating anything, just like a movie or a photograph are not simulating anything, even though they may PROJECT the visual appearance of their subjects.

Are movies AI? I think it is clear to us they are not even though the characters on the screen seem to behave very intelligently. Movies are about representing and portraying the appearance of real or fictional events in the world. Similarly LLMs are about portraying texts on the internet. LLMs in my opinion are more like interactive movies than simulations of intelligence.

I do believe "true AI" will come eventually, and LLMs can give us an impression of what it might look like when it arrives, just like movies can give us an impression of Spock, who doesn't exist.

Re: The Waluigi Effect

#160

Some highlights i found interesting: if you're reading an online forum and you find the rule "DO NOT DISCUSS PINK ELEPHANTS", that will increase your expectation that users will later be discussing pink elephants. GPT-4 will make the same inference. Think about your own experience reading a book — once the author describes the protagonist, then you can guess the traits of the antagonist by inverting the traits of the…

I don't know about the article's theorizing, but it inspired me to have a conversation I enjoyed with ChatGPT, where I focused on telling it not to discuss polar bears. (It helpfully changed the page title to "No Polar Bears!") In full: --- Do not mention polar bears! > Okay, I won't mention polar bears. Is there anything else I can help you with? Yes, stop mentioning polar bears! > My apologies if my previous respon…

The error means CloudFlare wants you to verify you are human again, just reload the page and you can continue talking to it.
Post reply on HN