Live data from Hacker News

Language models can explain neurons in language models

openai.com

81–90 of 497 posts

Re: Language models can explain neurons in language models

#81
post #68

Can anyone explain what they did? I’m not understanding from the webpage or the paper. What role does gpt4 play? I’m seeing they had gpt4 label every neuron but how?

I got the impression that it mentioned that the complexity of what's going on in GPT is so complex that we should use GPT to explain/summarize/graph what is going on.

We should ask AI, how are you doing this?

Re: Language models can explain neurons in language models

#82
post #68

Can anyone explain what they did? I’m not understanding from the webpage or the paper. What role does gpt4 play? I’m seeing they had gpt4 label every neuron but how?

I got the impression that it mentioned that the complexity of what's going on in GPT is so complex that we should use GPT to explain/summarize/graph what is going on. We should ask AI, how are you doing this?

Operator: Skynet, are you doing good thing? Skynet: Yes.

Re: Language models can explain neurons in language models

#83

Earlier quoted context omitted.

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

that is absolutely fascinating and also makes me extremely uncomfortable

Just wait until you notice how much humans do this day to day

Re: Language models can explain neurons in language models

#84

Earlier quoted context omitted.

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

Not supported by neuroimaging. Promoted without evidence or sufficient causal inference. https://www.health.harvard.edu/blog/right-brainleft-brain-ri... : > But, the evidence discounting the left/right brain concept is accumulating. According to a 2013 study from the University of Utah, brain scans demonstrate that activity is similar on both sides of the brain regardless of one's personality. > They looked at the br…

Agree I'm not a neuro subspecialist but I've listened to some talks at conferences out of interest and I don't think anyone still believes in this anymore. Anecdotally the few fMRI's I reported as a trainee didn't support this either.

Re: Language models can explain neurons in language models

#85

Earlier quoted context omitted.

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

that is absolutely fascinating and also makes me extremely uncomfortable

If that makes you uncomfortable you definitely should not go reading the evidence supporting the notion that conscious free will is an illusion.

https://www.mpg.de/research/unconscious-decisions-in-the-bra...

Re: Language models can explain neurons in language models

#86

Earlier quoted context omitted.

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

Not supported by neuroimaging. Promoted without evidence or sufficient causal inference. https://www.health.harvard.edu/blog/right-brainleft-brain-ri... : > But, the evidence discounting the left/right brain concept is accumulating. According to a 2013 study from the University of Utah, brain scans demonstrate that activity is similar on both sides of the brain regardless of one's personality. > They looked at the br…

You are talking about the popular narrative of “left brain” thinking being more logical and “right brain” thinking being more creative. You are correct this is unsupported.

The post you are replying to is talking about the small subset of individuals who have had their corpus callosum surgically severed, which makes it much more difficult for the brain to send messages between hemispheres. These patients exhibit “split brain” behavior that is well studied by experiments and can shed light into human consciousness and rationality.

Re: Language models can explain neurons in language models

#87
To me the value here is not that GPT4 has some special insight into explaining the behavior of GPT2 neurons (they say it's comparable to "human contractors" - but human performance on this task is also quite poor). The value is that you can just run this on every neuron if you're willing to spend the compute, and having a very fuzzy, flawed map of every neuron in a model is still pretty useful as a research tool.

But I would be very cautious about drawing conclusions from any individual neuron explanation generated in this way - even if it looks plausible by visual inspection of a few attention maps.

Re: Language models can explain neurons in language models

#88

Earlier quoted context omitted.

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

Not supported by neuroimaging. Promoted without evidence or sufficient causal inference. https://www.health.harvard.edu/blog/right-brainleft-brain-ri... : > But, the evidence discounting the left/right brain concept is accumulating. According to a 2013 study from the University of Utah, brain scans demonstrate that activity is similar on both sides of the brain regardless of one's personality. > They looked at the br…

Your response doesn't seem to be directly related to the previous poster's split-brain comments, but rather the popular misuse of the lateralization idea.

Re: Language models can explain neurons in language models

#89
Of note:

"... our technique works poorly for larger models, possibly because later layers are harder to explain."

And even for GPT-2, which is what they used for the paper:

"... the vast majority of our explanations score poorly ..."

Which is to say, we still have no clue as to what's going on inside GPT-4 or even GPT-3, which I think is the question many want an answer to. This may be the first step towards that, but as they also note, the technique is already very computationally intensive, and the focus on individual neurons as a function of input means that they can't "reverse engineer" larger structures composed of multiple neurons nor a neuron that has multiple roles; I would expect the former in particular to be much more common in larger models, which is perhaps why they're harder to analyze in this manner.

Re: Language models can explain neurons in language models

#90

Earlier quoted context omitted.

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

Not supported by neuroimaging. Promoted without evidence or sufficient causal inference. https://www.health.harvard.edu/blog/right-brainleft-brain-ri... : > But, the evidence discounting the left/right brain concept is accumulating. According to a 2013 study from the University of Utah, brain scans demonstrate that activity is similar on both sides of the brain regardless of one's personality. > They looked at the br…

This is not relevant to GP's comment. It has nothing to do with "are there fixed 'themes' that are operated in each hemisphere." It has to do with more generally, does the brain know what the brain is doing. The answer so far does not seem to be "yes."
Post reply on HN