Earlier quoted context omitted.
First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…
I'm not understanding the connection between your paragraphs here even after reading the first article. Even if you accept classic theory (e.g. hemispheric localization and the homunculus) which most experts don't all this suggests is that the brain tries to make sense of the information it has and in sparse environments it fills in. How does this make our behavior "mostly lies, fabrications, hallucinations, faulty r…
Language models can explain neurons in language models
91–100 of 497 posts
Re: Language models can explain neurons in language models
#92Earlier quoted context omitted.
> Why is everyone deferring to a bunch of effective altruism advocates when it comes to AI safety? I'm not sure Yudkowski is an EA, but the EAs want him in their polycule.
He posts on the forum. I'm not sure what more evidence is needed that he's part of it. https://forum.effectivealtruism.org/users/eliezeryudkowsky
Re: Language models can explain neurons in language models
#93Earlier quoted context omitted.
What if you ask it to emit the reflexive output, then feed that reflexive output back into the LLM for the conscious answer? What if you ask it to synthesize multiple internal streams of thought, for an ensemble of interior monologues, then have all those argue with each other using logic and then present a high level answer from that panoply of answers?
What if you do? LLMs don't have reflexive output or internal streams of thought, they are simply (complex) processes that produce streams of tokens based on an inputted stream of tokens. They don't have a special response to tokens that indicate higher-level thinking to humans.
Re: Language models can explain neurons in language models
#94> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.
You mean Yudkowski? I saw him on Lex Fridman and he was entirely unconvincing. Why is everyone deferring to a bunch of effective altruism advocates when it comes to AI safety?
Re: Language models can explain neurons in language models
#95> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.
You mean Yudkowski? I saw him on Lex Fridman and he was entirely unconvincing. Why is everyone deferring to a bunch of effective altruism advocates when it comes to AI safety?
Take this blog post for example, which between the lines reads: we don't expect to be able to align these systems ourselves, so instead we're hoping these systems are able to align each other.
Consider me not-very-soothed.
FWIW, there are plenty of AI experts who have been raising alarms as well. Hinton and Christiano, for example.
Re: Language models can explain neurons in language models
#96Earlier quoted context omitted.
The Goedel Incompleteness Theorem has no straightforward application to this question.
It would if the language model did reasoning according rules of logic. But they don't. They use Markov chains. To me it makes no sense to say that a LLM could explain its own reasoning if it does no (logical) reasoning at all. It might be able to explain how the neural network calculates its results. But there are no logical reasoning steps in there that could be explained, are there?
IANAE but although an LLM meets the definition of a Markov Chain as I understand it (current state in, probabilities of next states out), the big black box that spits out the probabilities could be doing anything.
Is it fundamentally impossible for reasoning to be an emergent property of an LLM, in a similar way to a brain? They can certainly do a good impression of logical reasoning- better than some humans in some cases?
Just because an LLM can be described as a Markov Chain doesn’t mean it _uses_ Markov Chains? An LLM is very different to the normal examples of Markov Chains I’m familiar with.
Or am I missing something?
In any case, coemu is an interesting related idea to constrain AIs to thinking in ways we can understand better:
https://futureoflife.org/podcast/connor-leahy-on-agi-and-cog...
https://www.alignmentforum.org/posts/ngEvKav9w57XrGQnb/cogni...
Re: Language models can explain neurons in language models
#97Earlier quoted context omitted.
You mean Yudkowski? I saw him on Lex Fridman and he was entirely unconvincing. Why is everyone deferring to a bunch of effective altruism advocates when it comes to AI safety?
I heard him on Lex too, and it seemed to be just a given that AI is going to be deceptive and want to kill us all. I don't think there was a single example of how that could be accomplished given. I'm open to hearing thoughts on this, maybe I'm not creative enough to see the 'obvious' ways this could happen.
I'm willing to bet the future of our species on my consistent victory in these types of matches, in fact.
Re: Language models can explain neurons in language models
#98I think this is a generous usage of "can." As the article admits, these explanations are 'imperfect' and I think that is definitely true.
It depends how you parse it. It is clearly true that they 'can' explain neurons, in the sense that at least some of the neurons are quite well explained. On the other hand, it's also the case that the vast majority of neurons are not well explained at all by this method (or likely any method). I̵t̵'̵s̵ ̵o̵n̵l̵y̵ ̵b̵e̵c̵a̵u̵s̵e̵ ̵o̵f̵ ̵a̵ ̵q̵u̵i̵r̵k̵ ̵o̵f̵ ̵A̵d̵a̵m̵W̵ ̵t̵h̵a̵t̵ ̵t̵h̵i̵s̵ ̵i̵s̵ ̵p̵o̵s̵s̵i̵b̵l̵e̵ ̵a̵t̵…
it does?
Re: Language models can explain neurons in language models
#99Earlier quoted context omitted.
that is absolutely fascinating and also makes me extremely uncomfortable
It's probably way worse than we can imagine. Reading/listening to someone like Robert Sapolsky [1] makes me laugh I could have ever hallucinated about such a muddy, not even wrong concept as "free will". Furthermore, between the brain and, say, the liver there is only a difference of speed/data integrity inasmuch as one cares to look for information processing as basal cognition: neurons firing in the brain, voltage-…
And that’s just the polar opposite of having a meaningful will at all. It is good that you are pretty much deterministic. You shouldn’t be deciding meaningful things randomly. If you made 20 copies of yourself and asked them to support or oppose some essential and important political question (about human rights, or war, or what-have-you) they should all come down on the same side. What kind of a Will would that be that chose randomly?
Re: Language models can explain neurons in language models
#100Earlier quoted context omitted.
First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…
I'm not understanding the connection between your paragraphs here even after reading the first article. Even if you accept classic theory (e.g. hemispheric localization and the homunculus) which most experts don't all this suggests is that the brain tries to make sense of the information it has and in sparse environments it fills in. How does this make our behavior "mostly lies, fabrications, hallucinations, faulty r…
But the bottom line is that introspection is not necessarily reliable.