Live data from Hacker News

Language models can explain neurons in language models

openai.com

71–80 of 497 posts

Re: Language models can explain neurons in language models

#71

Earlier quoted context omitted.

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

that is absolutely fascinating and also makes me extremely uncomfortable

This is why I suggest that curious individuals try a hallucinogen at least once*. It really makes the fragility of our perception, and how it’s held up mostly by itself, very apparent.

* in a safe setting with support, of course.

Re: Language models can explain neurons in language models

#72
post #2

> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.

DISCLAIMER: I think Yudkowsky is a serious thinker and his ideas should be taken seriously, regardless of whether not they are correct.

Your comment triggered a random thought: A perfect name for Yudkowsky et al and the AGI doomers is... wait for it... the Yuddites :)

Re: Language models can explain neurons in language models

#73
post #3

"This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself. A promising aspect of this approach is that it scales with the pace of AI development. As future models become increasingly intelligent and helpful as assistants, we will find better explanations." On first look this is genius but it seems pretty tautological in a way. How do we know i…

There is a longer-term problem of trusting the explainer system, but in the near-term that isn't really a concern.

The bigger value here in the near-term is _explicability_ rather than alignment per-se. Potentially having good explicability might provide insights into the design and architecture of LLMs in general, and that in-turn may enable better design of alignment-schemes.

Re: Language models can explain neurons in language models

#75
post #6

LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

I'm not understanding the connection between your paragraphs here even after reading the first article.

Even if you accept classic theory (e.g. hemispheric localization and the homunculus) which most experts don't all this suggests is that the brain tries to make sense of the information it has and in sparse environments it fills in.

How does this make our behavior "mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning" as most humans don't have a severed corpus callosum.

The discussion starts with:

"In a healthy human brain, these divergent hemispheric tendencies complement each other and create a balanced and flexible reasoning system. Working in unison, the left and right hemispheres can create inferences that have explanatory power and both internal and external consistency."

Re: Language models can explain neurons in language models

#76
post #60

Earlier quoted context omitted.

"LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own." There is no "their" and there is no "thought process" . There is something that produces text that appears to humans like there is something like thought going on (cf the Eliza Effect), but we must be wary of this anthropomorphising language. There is no self reflection, but if you ask an L…

What if you ask it to emit the reflexive output, then feed that reflexive output back into the LLM for the conscious answer? What if you ask it to synthesize multiple internal streams of thought, for an ensemble of interior monologues, then have all those argue with each other using logic and then present a high level answer from that panoply of answers?

What if you do? LLMs don't have reflexive output or internal streams of thought, they are simply (complex) processes that produce streams of tokens based on an inputted stream of tokens. They don't have a special response to tokens that indicate higher-level thinking to humans.

Re: Language models can explain neurons in language models

#77
post #6

LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

Not supported by neuroimaging. Promoted without evidence or sufficient causal inference.

https://www.health.harvard.edu/blog/right-brainleft-brain-ri... :

> But, the evidence discounting the left/right brain concept is accumulating. According to a 2013 study from the University of Utah, brain scans demonstrate that activity is similar on both sides of the brain regardless of one's personality.

> They looked at the brain scans of more than 1,000 young people between the ages of 7 and 29 and divided different areas of the brain into 7,000 regions to determine whether one side of the brain was more active or connected than the other side. No evidence of "sidedness" was found. The authors concluded that the notion of some people being more left-brained or right-brained is more a figure of speech than an anatomically accurate description.

Here's wikipedia on the topic: "Lateralization of brain function" https://en.wikipedia.org/wiki/Lateralization_of_brain_functi...

Furthermore, "Neuropsychoanalysis" https://en.wikipedia.org/wiki/Neuropsychoanalysis

Neuropsychology: https://en.wikipedia.org/wiki/Neuropsychology

Personality psychology > ~Biophysiological: https://en.wikipedia.org/wiki/Personality_psychology

MBTI > Criticism: https://en.wikipedia.org/wiki/Myers%E2%80%93Briggs_Type_Indi...

Connectome: https://en.wikipedia.org/wiki/Connectome

Re: Language models can explain neurons in language models

#78

Earlier quoted context omitted.

First of all, our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning: "In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw…

that is absolutely fascinating and also makes me extremely uncomfortable

Neurology is full of very uncomfortable facts. Here's one for you: there are patients who believe their arm is gone even though it's still there. When the doctor asks whose arm that is, they reply it must be someone else's. The brain can simply refuse to know something, and will adopt whatever delusions and contortions are necessary. Which of course leads to the realization that there could be things we're all incapable of knowing. There could be things right in front of our faces we simply refuse to perceive and we'd never know it.

Re: Language models can explain neurons in language models

#79
post #70
post #60

Earlier quoted context omitted.

"LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own." There is no "their" and there is no "thought process" . There is something that produces text that appears to humans like there is something like thought going on (cf the Eliza Effect), but we must be wary of this anthropomorphising language. There is no self reflection, but if you ask an L…

Or maybe the human thought process isn't as sophisticated as we imagined.

I'm not arguing for or against that. It's more the implications of sentience and selfhood implicit in the language many use around LLMs.

Re: Language models can explain neurons in language models

#80
post #61

Earlier quoted context omitted.

You mean Yudkowski? I saw him on Lex Fridman and he was entirely unconvincing. Why is everyone deferring to a bunch of effective altruism advocates when it comes to AI safety?

> Why is everyone deferring to a bunch of effective altruism advocates when it comes to AI safety? I'm not sure Yudkowski is an EA, but the EAs want him in their polycule.

He posts on the forum. I'm not sure what more evidence is needed that he's part of it.

https://forum.effectivealtruism.org/users/eliezeryudkowsky

Post reply on HN