Live data from Hacker News

Language models can explain neurons in language models

openai.com

251–260 of 497 posts

Re: Language models can explain neurons in language models

#252

Earlier quoted context omitted.

What's the argument that understanding neurons is necessary? Perhaps intelligence is like a black box input to our bodies (call it the "soul", even though this isn't testable and therefore not a hypothesis). The mind therefore wouldn't play any more of a role in intelligence than the eye. And I'm not sure people would say the eye is necessary for understanding intelligence. Now, I'm not really in a position to argue…

You can actually hypothesize that a soul exists and that intelligence is non-material, its just that your tests would quickly disprove that hypothesis - crude physical, mechanical modifications to the brain cause changes to intellect and character. If your hypothesis was correct you would not expect to see changes like that at all. Some people think that neurons specifically aren't necessary for understanding intelli…

I’m here playing devil’s advocate - this test doesn’t work. Here are some related thought experiments.

Suppose a soul is an immaterial source of intelligence, but it controls the body via machine-like material hardware such as neurons.

Or an alternative, suppose there is a soul inside your body “watching” the sensory activations within your brain like a movie. The brain and body create the movie & have some intelligence, but other important properties of the consciousness are bound to this observer entity.

In both these cases, the test just shows that if you damage the hardware, you can no longer observe intelligence because you’ve broken the end-to-end flow of the machine.

Re: Language models can explain neurons in language models

#253
post #85

Earlier quoted context omitted.

If that makes you uncomfortable you definitely should not go reading the evidence supporting the notion that conscious free will is an illusion. https://www.mpg.de/research/unconscious-decisions-in-the-bra...

i always thought that the concept of free will didnt bother me. but it turns out i just didn’t understand what it implied. oh dear. if it turns out that true, it’s truly amazing how well we convince ourselves that we’re in control. but if our brain controls our actions and not our consciousness, then what is the purpose of consciousness?

Perhaps there is no purpose to consciousness.

Perhaps it's a phenomenon that somehow arises independently ex nihilo from sufficiently complex systems, only ever able to observe, unable to act.

Weird to think about.

Re: Language models can explain neurons in language models

#254

Earlier quoted context omitted.

Evolution created intelligence without even being intelligent itself

How do you know that’s true?

Evolution is just a description of a process, it isn't a tangible thing.

Re: Language models can explain neurons in language models

#255
post #62

This isnt exactly building an understanding of LLMs from first principles... IMO we should broadly be following the (imperfect) example set forth by neuroscientists attempting to explain fMRI scans and assigning functionality to various subregions in the brain. It is circular and "unsafe" from an alignment perspective to use a complex model to understand the internals of a simpler model; in order to understand GPT4 t…

I've been working in systems neuroscience for a few years (something of a combination lab tech/student, so full disclosure, not an actual expert). Based on my experience with model organisms (flies & rats, primarily), it is actually pretty amazing how analogous the techniques and goals used in this sort of research are to those we use in systems neuroscience. At a very basic level, the primary task of correlating neu…

Yup, this article on predictive coding for example is particularly interesting.

Lots of parallels to how our brains are thought to work.

https://en.m.wikipedia.org/wiki/Predictive_coding

Re: Language models can explain neurons in language models

#256
post #89

Of note: "... our technique works poorly for larger models, possibly because later layers are harder to explain." And even for GPT-2, which is what they used for the paper: "... the vast majority of our explanations score poorly ..." Which is to say, we still have no clue as to what's going on inside GPT-4 or even GPT-3, which I think is the question many want an answer to. This may be the first step towards that, bu…

Funny that we never quite understood how intelligence worked and yet it appears that we're pretty damn close to recreating it - still without knowing how it works. I wonder how often this happens in the universe...

Starting a fire is easy to do even if you don't know how it works.

Re: Language models can explain neurons in language models

#257

Earlier quoted context omitted.

Because we eat and breath through the same tube

There are two tubes.

Evolution came up with the shared eating/breathing tube design because it made sense for aquatic animals (from which we evolved).

Re: Language models can explain neurons in language models

#258

Earlier quoted context omitted.

If you spent even more time with GPT-4 it would be evident that it is definitely not. Especially if you try to use it as some kind of autonomous agent.

If you spent even more time with GPT-4 it would be evident that it definitely is. Especialy if you try to use it as some kind of autonomous agent. (Notice how baseless comments can sway either way)

> (Notice how baseless comments can sway either way)

No they can’t! ;)

Re: Language models can explain neurons in language models

#259
post #219

Earlier quoted context omitted.

> that we're pretty damn close to recreating it Is that evident already or are we fitting the definition of intelligence without being aware?

If you spent any time with GPT-4 it should be evident.

Compare gpt-4 with a baby and you'll see that predicting the next word in sequence is not human intelligence

Re: Language models can explain neurons in language models

#260
Commenters here seem a little fixated on the fact that the technique scores poorly. This is true and somewhat problematic but exploring the data it looks to me like this could be more a problem with a combination of the scoring function and the limits of language rather than the methodology.

For example, look at https://openaipublic.blob.core.windows.net/neuron-explainer/...

It's described as "expressions of completion or success" with a score of 0.38. But going through the examples, they are very consistently a sort of colloquial expression of "completion/success" with a touch of surprise and maybe challenge.

Examples are like: "Nuff said", "voila!", "Mission accomplished", "Game on!", "End of story", "enough said", "nailed it" etc.

If they expressed it as a basket of words instead of a sentence, and could come up with words that express it better I'd score it much higher.

Post reply on HN