Live data from Hacker News

Language models can explain neurons in language models

openai.com

61–70 of 497 posts

Re: Language models can explain neurons in language models

#61
post #2

> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.

You mean Yudkowski? I saw him on Lex Fridman and he was entirely unconvincing. Why is everyone deferring to a bunch of effective altruism advocates when it comes to AI safety?

> Why is everyone deferring to a bunch of effective altruism advocates when it comes to AI safety?

I'm not sure Yudkowski is an EA, but the EAs want him in their polycule.

Re: Language models can explain neurons in language models

#62

This isnt exactly building an understanding of LLMs from first principles... IMO we should broadly be following the (imperfect) example set forth by neuroscientists attempting to explain fMRI scans and assigning functionality to various subregions in the brain. It is circular and "unsafe" from an alignment perspective to use a complex model to understand the internals of a simpler model; in order to understand GPT4 t…

I've been working in systems neuroscience for a few years (something of a combination lab tech/student, so full disclosure, not an actual expert).

Based on my experience with model organisms (flies & rats, primarily), it is actually pretty amazing how analogous the techniques and goals used in this sort of research are to those we use in systems neuroscience. At a very basic level, the primary task of correlating neuron activation to a given behavior is exactly the same. However, ML researchers benefit from data being trivial to generate and entire brains being analyzable in one shot as a result, whereas in animal research elucidating the role of neurons in a single circuit costs millions of dollars and many researcher-years.

The similarities between the two are so clear that I noticed that in its Microscope tool [1], OpenAI even refers to the models they are studying as "model organisms", an anthropomorphization which I find very apt. Another article I saw a while back on HN which I thought was very cool was [2], which describes the task of identifying the role of a neuron responsible for a particular token of output. This one is especially analogous because it operates on such a small scale, much closer to what systems neuroscientists studying model organisms do.

[1] https://openai.com/research/microscope [2] https://clementneo.com/posts/2023/02/11/we-found-an-neuron

Re: Language models can explain neurons in language models

#63
post #3

"This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself. A promising aspect of this approach is that it scales with the pace of AI development. As future models become increasingly intelligent and helpful as assistants, we will find better explanations." On first look this is genius but it seems pretty tautological in a way. How do we know i…

Why is this genius? It's just the NN equivalent of making a new programming language and getting it to the point where its compiler can be written in itself.

The reliability question is of course the main issue. If you don't know how the system works, you can't assign a trust value to anything it comes up with, even if it seems like what it comes up with makes sense.

Re: Language models can explain neurons in language models

#64
post #60
post #6

LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…

"LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own." There is no "their" and there is no "thought process" . There is something that produces text that appears to humans like there is something like thought going on (cf the Eliza Effect), but we must be wary of this anthropomorphising language. There is no self reflection, but if you ask an L…

What if you ask it to emit the reflexive output, then feed that reflexive output back into the LLM for the conscious answer?

What if you ask it to synthesize multiple internal streams of thought, for an ensemble of interior monologues, then have all those argue with each other using logic and then present a high level answer from that panoply of answers?

Re: Language models can explain neurons in language models

#65

Earlier quoted context omitted.

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

that is absolutely fascinating and also makes me extremely uncomfortable

Same. We are so, so profoundly not what it feels like we are, to most of us anyway.

I am morbidly curious how people are going to creatively explain away the more challenging insights AI gives us in to what consciousness is.

Re: Language models can explain neurons in language models

#66
post #6

LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

Phantoms in the Brain is a fascinating book that deals with exactly this topic.

Re: Language models can explain neurons in language models

#67

Earlier quoted context omitted.

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

that is absolutely fascinating and also makes me extremely uncomfortable

It's probably way worse than we can imagine.

Reading/listening to someone like Robert Sapolsky [1] makes me laugh I could have ever hallucinated about such a muddy, not even wrong concept as "free will".

Furthermore, between the brain and, say, the liver there is only a difference of speed/data integrity inasmuch as one cares to look for information processing as basal cognition: neurons firing in the brain, voltage-gated ion channels and gap junctions controlling bioelectrical gradients in the liver, and almost everywhere in the body. Why does only the brain has a "feels like" sensation? The liver may have one as well, but the brain being an autarchic dictator perhaps suppresses the feeling of the liver, it certainly abstracts away the thousands of highly specialized decisions the liver takes each second solving adequately the complex problem space of blood processing. Perhaps Thomas Nagel shouldn't have asked "What Is It Like to Be a Bat?" [2] but what is it like to be a liver.

[1] "Robert Sapolsky: Justice and morality in the absence of free will", https://www.youtube.com/watch?v=nhvAAvwS-UA

[2] https://en.wikipedia.org/wiki/What_Is_It_Like_to_Be_a_Bat%3F

Re: Language models can explain neurons in language models

#69
post #11
post #4

Has anyone here found a link to the actual paper? If I click on 'paper', I only see what seems to be an awkward HTML version.

You mean this? https://openaipublic.blob.core.windows.net/neuron-explainer/... Would you prefer a PDF? (I'm always fascinated to hear from people who would rather read a PDF than a web-native paper like this one, especially given that web papers are actually readable on mobile devices. Do you do all of your reading on a laptop?)

> always fascinated

That feels like a loaded phrase. Is it "false confusion" adjacent?

Re: Language models can explain neurons in language models

#70
post #60
post #6

LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…

"LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own." There is no "their" and there is no "thought process" . There is something that produces text that appears to humans like there is something like thought going on (cf the Eliza Effect), but we must be wary of this anthropomorphising language. There is no self reflection, but if you ask an L…

Or maybe the human thought process isn't as sophisticated as we imagined.
Post reply on HN