Live data from Hacker News

Language models can explain neurons in language models

openai.com

121–130 of 497 posts

Re: Language models can explain neurons in language models

#121

Earlier quoted context omitted.

This is also why I go into chess matches against 1400 elo players. I cannot conceive of the specific ways in which they will beat me (a 600 elo player), so I have good reason to suspect that I can win. I'm willing to bet the future of our species on my consistent victory in these types of matches, in fact.

Again, a given that AI is adversarial. Edit: In addition, as an 1100 elo chess player, I can very easily tell you how a 1600 player is going to beat me. The analogy doesn't hold. I'm in good faith asking how AI could destroy humanity. It seems given the confidence people who are scared of AI have in this, that they have some concrete examples in mind.

No it’s a given that some people who attempt to wield AI will be adversarial.

In any case a similar argument can be made with merely instrumental goals causing harm: “I am an ant and I do not see how or why a human would cause me harm, therefore I am not in danger.”

Re: Language models can explain neurons in language models

#122

Earlier quoted context omitted.

You mean Yudkowski? I saw him on Lex Fridman and he was entirely unconvincing. Why is everyone deferring to a bunch of effective altruism advocates when it comes to AI safety?

I heard him on Lex too, and it seemed to be just a given that AI is going to be deceptive and want to kill us all. I don't think there was a single example of how that could be accomplished given. I'm open to hearing thoughts on this, maybe I'm not creative enough to see the 'obvious' ways this could happen.

IMHO the argument isn't that AI is definitely going to be deceptive and want to kill us all, but rather that if you're 90% sure that AI is going to be just fine, that 10% of existential risk is simply not acceptable, so you should assume that this level of certainty isn't enough and you should act as if AI may be deceptive and may kill us all and take very serious preventive measures even if you're quite certain that it won't be needed - because "quite certain" isn't enough, you want to be at "this is definitely established to not lead to Skynet" level.

Re: Language models can explain neurons in language models

#123

Earlier quoted context omitted.

I heard him on Lex too, and it seemed to be just a given that AI is going to be deceptive and want to kill us all. I don't think there was a single example of how that could be accomplished given. I'm open to hearing thoughts on this, maybe I'm not creative enough to see the 'obvious' ways this could happen.

This is also why I go into chess matches against 1400 elo players. I cannot conceive of the specific ways in which they will beat me (a 600 elo player), so I have good reason to suspect that I can win. I'm willing to bet the future of our species on my consistent victory in these types of matches, in fact.

Other caveman use fire to cook food. Fire scary and hurt. No understand fire. Fire cavemen bad.

Re: Language models can explain neurons in language models

#124

Earlier quoted context omitted.

He didn't say behavior. He said explanations of behaviour. Split brain experiments aside, this is pretty evident from other research. We can't recreate previous mental states, we just do a pretty good job (usually) of rationalizing decisions after the fact. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3196841/

I'm reading this as our explanations for our own behavior as in why am I typing on this keyboard right now, in which case it's not evident at all. The existence of cognitive dissonance suggested in your citation is in no way analogous to "our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning" and in fact supports the opposite.

One primary explanation of ourselves is that there is in fact a "self" there, we feel this "self" as being permanent, continuous through time, yet we are absolutely sure that is a lie: there are no continuous processes in the entire universe, energy itself is quantized.

In the morning when we wake up, we are "booting" up the memories the brain finds and we believe that we have persisted through time, from yesterday to today, yet we are absolutely sure that is a lie: just look at an Alzheimer patient.

We are feeling this self as if it's somewhere above the neck and we feel like this self is looking at the world and sees "out there", yet we are absolutely sure that is a lie: our senses are being overflown by inputs and the brain filters them, shapes a model of the world, and presents that model to the internal model of itself, which gets so immersed into model of the world that starts to believe the model is indeed the world, until the first bistable image [1] breaks the model down.

[1] https://www.researchgate.net/profile/Amanda-Parker-14/public...

Re: Language models can explain neurons in language models

#125

Earlier quoted context omitted.

This is not relevant to GP's comment. It has nothing to do with "are there fixed 'themes' that are operated in each hemisphere." It has to do with more generally, does the brain know what the brain is doing. The answer so far does not seem to be "yes."

Says who? There is actual evidence to support that our brain doesn't "know" what it is doing on a subconscious level? As far as I'm aware it's more that conscious humans don't understand how our brain works. I think the correct statement is "so far the answer is we don't know"

The split brain experiments very very clearly indicate that different parts of the brain can independently conduct behavior and gain knowledge independently of other parts.

How or if this generalizes to healthy brains is not super clear, but it does actually provide a good explanatory model for all sorts of self-contradictory behavior (like addiction): the brain has many semi-independent “interests” that are jockeying for overall control of the organism’s behavior. These interests can be fully contradictory to each other.

Correct, ultimately we do not know. But it’s actually a different question than your rephrasing.

Re: Language models can explain neurons in language models

#126

Earlier quoted context omitted.

Again, a given that AI is adversarial. Edit: In addition, as an 1100 elo chess player, I can very easily tell you how a 1600 player is going to beat me. The analogy doesn't hold. I'm in good faith asking how AI could destroy humanity. It seems given the confidence people who are scared of AI have in this, that they have some concrete examples in mind.

No it’s a given that some people who attempt to wield AI will be adversarial. In any case a similar argument can be made with merely instrumental goals causing harm: “I am an ant and I do not see how or why a human would cause me harm, therefore I am not in danger.”

People wielding AI and destroying humanity is very different from AI itself, being a weird alien intelligence, destroying humanity.

Honestly if you have no examples you can't really blame people for not being scared. I have no reason to think this ant-human relationship is analogous.

And seriously, I've made no claims that AI is benign so please stop characterizing my claims thusly. The question is simple, give me a single hypothetical example of how an AI will destroy humanity?

Re: Language models can explain neurons in language models

#127
post #60

Earlier quoted context omitted.

"LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own." There is no "their" and there is no "thought process" . There is something that produces text that appears to humans like there is something like thought going on (cf the Eliza Effect), but we must be wary of this anthropomorphising language. There is no self reflection, but if you ask an L…

As we don't know for sure what is happening 100% within a neural network, we can say we don't believe that they're thinking and we would still need to define the word thinking. Once LLM's can self-modify, the word "thinking" will be more accurate than it is today. And when Hinton says at MIT, "I find it very hard to believe that they don't have semantics when they consult problems like you know how I paint the rooms…

In this case, I think we do if you will check out the paper (https://openaipublic.blob.core.windows.net/neuron-explainer/...). Their method is to

1. Show GPT-4 a GPT-produced text with the activation level of a specific neuron at the time it was producing that part of the text highlighted. They then ask GPT-4 for an explanation of what the neuron is doing.

Text: "...mathematics is done _properly_, it...if it's done _right_. (Take ..."

GPT produces "words and phrases related to performing actions correctly or properly".

2. Based on the explanation, get GPT to guess how strong the neuron activates on a new text.

"Assuming that the neuron activates on words and phrases related to performing actions correctly or properly. GPT-4 guesses how strongly the neuron responds at each token: '...Boot. When done _correctly_, "Secure...'"

3. Compare those predictions to the actual activations of the neuron on the text to generate a score.

So there is no introspection going on.

They say, "We applied our method to all MLP neurons in GPT-2 XL [out of 1.5B?]. We found over 1,000 neurons with explanations that scored at least 0.8, meaning that according to GPT-4 they account for most of the neuron's top-activating behavior." But they also mention, "However, we found that both GPT-4-based and human contractor explanations still score poorly in absolute terms. When looking at neurons, we also found the typical neuron appeared quite polysemantic."

Re: Language models can explain neurons in language models

#128
post #60
post #6

LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…

"LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own." There is no "their" and there is no "thought process" . There is something that produces text that appears to humans like there is something like thought going on (cf the Eliza Effect), but we must be wary of this anthropomorphising language. There is no self reflection, but if you ask an L…

Pride comes before the fall, and the AI comes before humility.

Re: Language models can explain neurons in language models

#129
post #103

Earlier quoted context omitted.

People won't care until an actually scary AI exists. Will be easy to stop at that point. Or you can just stop research here and hope another country doesn't get one first. Im personally skeptical it will exist. Honestly might be making it worse with the scaremongering coming from uncharismatic AI alignment people.

Why would it be easy to stop at that point? The believable value prop will increase in lockstep with the believable scare factor, not to mention the (already significant) proliferation out of ultra expensive research orgs into open source repos. Nuclear weapons proliferated explicitly because they proved their scariness.

If AI can exist humans have to figure it out. It’s what we do. Really shockingly delusional to think people are gonna use chatgpt for a few min get bored and then ban it like it’s a nuke. I’d rather the USA get it first anyways.

Re: Language models can explain neurons in language models

#130
post #87

To me the value here is not that GPT4 has some special insight into explaining the behavior of GPT2 neurons (they say it's comparable to "human contractors" - but human performance on this task is also quite poor). The value is that you can just run this on every neuron if you're willing to spend the compute, and having a very fuzzy, flawed map of every neuron in a model is still pretty useful as a research tool. But…

They also mention they got a score above 0.8 for 1000 neurons out of GPT2 (which has 1.5B (?)).
Post reply on HN