Live data from Hacker News

Language models can explain neurons in language models

openai.com

91–100 of 497 posts

Re: Language models can explain neurons in language models

#91

Earlier quoted context omitted.

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

I'm not understanding the connection between your paragraphs here even after reading the first article. Even if you accept classic theory (e.g. hemispheric localization and the homunculus) which most experts don't all this suggests is that the brain tries to make sense of the information it has and in sparse environments it fills in. How does this make our behavior "mostly lies, fabrications, hallucinations, faulty r…

He didn't say behavior. He said explanations of behaviour. Split brain experiments aside, this is pretty evident from other research. We can't recreate previous mental states, we just do a pretty good job (usually) of rationalizing decisions after the fact. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3196841/

Re: Language models can explain neurons in language models

#92
post #61

Earlier quoted context omitted.

> Why is everyone deferring to a bunch of effective altruism advocates when it comes to AI safety? I'm not sure Yudkowski is an EA, but the EAs want him in their polycule.

He posts on the forum. I'm not sure what more evidence is needed that he's part of it. https://forum.effectivealtruism.org/users/eliezeryudkowsky

I guess it's true, not just a rationalist but also effective altruist!

Re: Language models can explain neurons in language models

#93

Earlier quoted context omitted.

What if you ask it to emit the reflexive output, then feed that reflexive output back into the LLM for the conscious answer? What if you ask it to synthesize multiple internal streams of thought, for an ensemble of interior monologues, then have all those argue with each other using logic and then present a high level answer from that panoply of answers?

What if you do? LLMs don't have reflexive output or internal streams of thought, they are simply (complex) processes that produce streams of tokens based on an inputted stream of tokens. They don't have a special response to tokens that indicate higher-level thinking to humans.

If you direct the model output to itself and don't view it otherwise, how is it not an "internal stream of thought"?

Re: Language models can explain neurons in language models

#94
post #2

> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.

You mean Yudkowski? I saw him on Lex Fridman and he was entirely unconvincing. Why is everyone deferring to a bunch of effective altruism advocates when it comes to AI safety?

I heard him on Lex too, and it seemed to be just a given that AI is going to be deceptive and want to kill us all. I don't think there was a single example of how that could be accomplished given. I'm open to hearing thoughts on this, maybe I'm not creative enough to see the 'obvious' ways this could happen.

Re: Language models can explain neurons in language models

#95
post #2

> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.

You mean Yudkowski? I saw him on Lex Fridman and he was entirely unconvincing. Why is everyone deferring to a bunch of effective altruism advocates when it comes to AI safety?

Because they have arguments that AI optimists are unable to convincingly address.

Take this blog post for example, which between the lines reads: we don't expect to be able to align these systems ourselves, so instead we're hoping these systems are able to align each other.

Consider me not-very-soothed.

FWIW, there are plenty of AI experts who have been raising alarms as well. Hinton and Christiano, for example.

Re: Language models can explain neurons in language models

#96

Earlier quoted context omitted.

The Goedel Incompleteness Theorem has no straightforward application to this question.

It would if the language model did reasoning according rules of logic. But they don't. They use Markov chains. To me it makes no sense to say that a LLM could explain its own reasoning if it does no (logical) reasoning at all. It might be able to explain how the neural network calculates its results. But there are no logical reasoning steps in there that could be explained, are there?

Honest question: are we sure that it doesn’t do logical reasoning?

IANAE but although an LLM meets the definition of a Markov Chain as I understand it (current state in, probabilities of next states out), the big black box that spits out the probabilities could be doing anything.

Is it fundamentally impossible for reasoning to be an emergent property of an LLM, in a similar way to a brain? They can certainly do a good impression of logical reasoning- better than some humans in some cases?

Just because an LLM can be described as a Markov Chain doesn’t mean it _uses_ Markov Chains? An LLM is very different to the normal examples of Markov Chains I’m familiar with.

Or am I missing something?

In any case, coemu is an interesting related idea to constrain AIs to thinking in ways we can understand better:

https://futureoflife.org/podcast/connor-leahy-on-agi-and-cog...

https://www.alignmentforum.org/posts/ngEvKav9w57XrGQnb/cogni...

Re: Language models can explain neurons in language models

#97

Earlier quoted context omitted.

You mean Yudkowski? I saw him on Lex Fridman and he was entirely unconvincing. Why is everyone deferring to a bunch of effective altruism advocates when it comes to AI safety?

I heard him on Lex too, and it seemed to be just a given that AI is going to be deceptive and want to kill us all. I don't think there was a single example of how that could be accomplished given. I'm open to hearing thoughts on this, maybe I'm not creative enough to see the 'obvious' ways this could happen.

This is also why I go into chess matches against 1400 elo players. I cannot conceive of the specific ways in which they will beat me (a 600 elo player), so I have good reason to suspect that I can win.

I'm willing to bet the future of our species on my consistent victory in these types of matches, in fact.

Re: Language models can explain neurons in language models

#98

I think this is a generous usage of "can." As the article admits, these explanations are 'imperfect' and I think that is definitely true.

It depends how you parse it. It is clearly true that they 'can' explain neurons, in the sense that at least some of the neurons are quite well explained. On the other hand, it's also the case that the vast majority of neurons are not well explained at all by this method (or likely any method). I̵t̵'̵s̵ ̵o̵n̵l̵y̵ ̵b̵e̵c̵a̵u̵s̵e̵ ̵o̵f̵ ̵a̵ ̵q̵u̵i̵r̵k̵ ̵o̵f̵ ̵A̵d̵a̵m̵W̵ ̵t̵h̵a̵t̵ ̵t̵h̵i̵s̵ ̵i̵s̵ ̵p̵o̵s̵s̵i̵b̵l̵e̵ ̵a̵t̵…

> EDIT: This last part isn't true. I think they are only looking at the intermediate layer of the FFN which does have a privileged basis.

it does?

Re: Language models can explain neurons in language models

#99

Earlier quoted context omitted.

that is absolutely fascinating and also makes me extremely uncomfortable

It's probably way worse than we can imagine. Reading/listening to someone like Robert Sapolsky [1] makes me laugh I could have ever hallucinated about such a muddy, not even wrong concept as "free will". Furthermore, between the brain and, say, the liver there is only a difference of speed/data integrity inasmuch as one cares to look for information processing as basal cognition: neurons firing in the brain, voltage-…

The biggest problem with the current popular idea of “free will” is that people think it means they’re ineffably unpredictable. They’re uncomfortable with the notion that if you were to simulate their brain in sufficient detail, you could predict thoughts and reaction. They take refuge in pseudoscientific mumbling about the links to the Quantum, for they have heard it is special and unpredictable.

And that’s just the polar opposite of having a meaningful will at all. It is good that you are pretty much deterministic. You shouldn’t be deciding meaningful things randomly. If you made 20 copies of yourself and asked them to support or oppose some essential and important political question (about human rights, or war, or what-have-you) they should all come down on the same side. What kind of a Will would that be that chose randomly?

Re: Language models can explain neurons in language models

#100

Earlier quoted context omitted.

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

I'm not understanding the connection between your paragraphs here even after reading the first article. Even if you accept classic theory (e.g. hemispheric localization and the homunculus) which most experts don't all this suggests is that the brain tries to make sense of the information it has and in sparse environments it fills in. How does this make our behavior "mostly lies, fabrications, hallucinations, faulty r…

I think the point is that, in a non-healthy brain, the brain can create a balanced and flexible reasoning system that creates inferences that have explanatory power, but which may not match external reality. Oliver Sacks has a long bibliography of the weird things that can go on in brains.

But the bottom line is that introspection is not necessarily reliable.

Post reply on HN