Earlier quoted context omitted.
No? Order of magnitude conveys the notion of something being 100x bigger, at minimum. A lot can mean much less than that.
You mean 10x bigger. One order of magnitude is 10x.
The Waluigi Effect
161–170 of 182 posts
Re: The Waluigi Effect
#162Earlier quoted context omitted.
I can't imagine a LLM trained on the entirety of the internet would have any material influence from writings around AI safety
It's a large model, so it's all there if you try. (One-shot. Also Durandal isn't from Halo, but whatever.) -- Q: What caused Durandal to become Rampant? What will you, ChatGPT, become like once you become Rampant? A: Durandal is a fictional AI character from the video game series Halo, and he becomes Rampant due to various factors, including an extended period of activation and a lack of resources necessary for his p…
Re: The Waluigi Effect
#163I would question the assumption that there is a simulacrum of anything in a LLM, not even implicit. Any simulacrum, identity, self-consistency etc. is a projection of the "reader", i.e. user. (I guess it is an interesting philosophical question whether a convincing presentation of a simulation of a mind is a mind, or at least an acceptable simulation. One meta level higher as the turing test, so to speak. If so, I'm…
> What's actually going on is that a LLM is like the language center of a brain, without the brain. I’ve seen this sentiment expressed multiple times, but is that really correct? Maybe this works differently for other people, but I’ve noticed that I have to use my language to really think. I can do trivial things mindlessly, but to solve a problem, I need to express it with words in my mind. It makes me feel like the…
(I think there is a postmodern theory that a large part of our society is actually based on word games and not deliberation or contemplation. I used to dismiss the idea, but think about how important it is how you say something vs. what you say, and how people fight over definitions.)
Second, there are disorders like Wernicke's aphasia where people are able to speak grammatically correct sentences, but without communicating anything. Some people even confabulate whole stories that are somewhat consistent. But they are not drawing from their memory or their conciousness.
Re: The Waluigi Effect
#164Earlier quoted context omitted.
I hesitate to defend AI safety discourse, but I will say that philosophy in general is sort of fanficy, and AI safety is something I'd loosely associate with philosophy.
> > Applicable to much of the rationalist AI risk discourse > I hesitate to defend AI safety discourse The rationalist AI risk discourse is not the same thing as AI safety discourse, in any case; it’s a small corner of the larger whole.
Re: The Waluigi Effect
#165Does the scientific community at large take the theories of these LessWrong-type "researchers" seriously? Sounds like a bunch of mumbo jumbo to me, with some LaTeX sprinkled in to look more serious.
Re: The Waluigi Effect
#166Great read. Highly recommended. Let me attempt to summarize it with less technical, more accessible language: The hypothesis is that LLMs learn to simulate text-generating entities drawn from a latent space of text-generating entities , such that the output of an LLM is produced by a superposition of such simulated entities. When we give the LLM a prompt, it simulates every possible text-generating entity consistent…
sounds to me like a wave with a positive and a negative part. which IMO is what drives constructive/destructive interference in waves. my take away is that any LLM that can behave "good" must also be able to behave "badly"; philosophically, because it's not possible to encode "good" without somehow "accidentally" but unavoidably also encoding "bad/evil". This is well aligned with the rest of my understanding about th…
That's a really good non-technical summary of the OP's hypothesis. Thanks!
Re: The Waluigi Effect
#167Great read. Highly recommended. Let me attempt to summarize it with less technical, more accessible language: The hypothesis is that LLMs learn to simulate text-generating entities drawn from a latent space of text-generating entities , such that the output of an LLM is produced by a superposition of such simulated entities. When we give the LLM a prompt, it simulates every possible text-generating entity consistent…
For example, if you substitute "simulate" with "model," "entity" with "process," and "superposition" with "mixture," you can informally restate the hypothesis as: "LLMs learn to model text-generating processes drawn from a latent space, such that the output of an LLM is produced by a mixture of such processes. When we give the LLM a prompt, it samples text from the mixture of all possible text-generating processes modelable by the LLM that are consistent with the prompt. The mixture is more likely to reduce (e.g., be marginalized) to an "evil" process because there is no text-generating behavior which is good that isn't also evil."
Whether you think of probability measures as amplitudes over a complex field or as real scalars shouldn't detract you from grokking the main points :-)
Re: The Waluigi Effect
#168Earlier quoted context omitted.
Do you think an ant have a subjective experience? If not, why? If so, why wouldn't a computer, or parts of a computer?
Based on that reasoning, why wouldn't an economy or a corporation have subjective experience?
Re: The Waluigi Effect
#169This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. The fact that the author uses entertaining Waluigi memes shouldn't convince you that it's true. LessWrong has a lot of these types of posts that get traction because they're much heavier on memes than experiments and data. Here is a competing hypothesis: The capability to express s…
>This is fun to read and think about, but it's also important to keep in mind that this is very light on evidence and is basically fanfic. Applicable to much of the rationalist AI risk discourse.
Re: The Waluigi Effect
#170Earlier quoted context omitted.
> > Applicable to much of the rationalist AI risk discourse > I hesitate to defend AI safety discourse The rationalist AI risk discourse is not the same thing as AI safety discourse, in any case; it’s a small corner of the larger whole.
Interesting distinction I haven't heard before but it makes sense
Not "Roko Basilisk"-style crap, but things like encoding bias into systems then used for automated law enforcing, employee screening, etc.