Live data from Hacker News

Language models can explain neurons in language models

openai.com

261–270 of 497 posts

Re: Language models can explain neurons in language models

#261

Earlier quoted context omitted.

> It is also the reason why they cannot be trusted in the most serious of applications which such decision making requires lots of transparency rather than a model regurgitating nonsense confidently. Doesn't this criticism also apply to people to some extent? We don't know what the purpose of individual brain neurons is.

As a person I can at least tell you what I do and don't understand about something, ask questions to improve/correct my understanding, and truthfully explain my perspective and reasoning. The machine model is not only a black box, but one incapable of understanding anything about its input, "thought process", or output. It will blindly spit out a response based on its training data and weights, without knowing the di…

As a person, you can tell what you think you do and don't understand, and you can explain what you think your reasoning is. In practice, people get both wrong all the time. People aren't always truthful about it, either, and there's no reliable way to tell if they are.

Re: Language models can explain neurons in language models

#262
post #207

I built a toy neural network that runs in the browser[1] to model 2D functions with the goal of doing something similar to this research (in a much more limited manner, ofc). Since the input space is so much more limited than language models or similar, it's possible to examine the outputs for each neuron for all possible inputs, and in a continuous manner. In some cases, you can clearly see neurons that specialize t…

This tool is really lovely, great work!

I'd be curious to see Softmax Linear Units [1] integrated into the possible activation functions since they seem to improve interpretability.

PS: I share your curiosity with respect to things like deep dream. My brief summary of this paper is that you can use GPT4 to summarize what's similar about a set of highlighted words in context which is clever but doesn't fundamentally inform much that we didn't already know about how these models work. I wonder if there's some diffusion based approach that could be used to diffuse from noise in the residual stream towards a maximized activation at a particular point.

[1] https://transformer-circuits.pub/2022/solu/index.html

Re: Language models can explain neurons in language models

#263

Earlier quoted context omitted.

I have to agree. I often think, “maybe I should use ChatGPT for this” then I realise I have very little way to verify what it tells me and as someone working in engineering, If I don’t understand the black box, I just can’t do it. I’m attracted to open source, because I can look at the code understand it.

Open source doesn't mean you can explain the black box any better. and humans are black boxes that don't understand their mental processes either. We're currently better than LLMs at it i suppose but we're still very poor at it. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3196841/ https://pure.uva.nl/ws/files/25987577/Split_Brain.pdf https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4204522 We can't recreate previous men…

Humans can be held accountable so it’s not the same. Even if we’re a black box, we share common traits with other humans. We’re trained in similar ways. So we mostly understand what we will and won’t do.

I think this constant degradation of humans is really foolish and harmful personally. “We’re just black boxes etc”, we might not know how brains work but we do and can understand each other.

On the other hand I’m starting to feel like “AI researchers” are the greatest black box I’ve ever seen, the logic of what they’re trying to create and their hopes for it really baffle me.

By the way, I have infinitely more hope of understanding an open source black box compared to a closed source one?

Re: Language models can explain neurons in language models

#264
post #219

Earlier quoted context omitted.

If you spent any time with GPT-4 it should be evident.

If you spent even more time with GPT-4 it would be evident that it is definitely not. Especially if you try to use it as some kind of autonomous agent.

AI research has put hardly any effort into building goal-directed agents / A-Life since the advent of Machine Learning. A-Life was last really "looked into" in the '70s, back when "AI" meant Expert Systems and Behavior Trees.

All the effort in AI research since the advent of Machine Learning, has been focused on making systems that — in neurological terms — are given a sensory stimulus of a question, and then passively "dream" a response to said question as a kind of autonomic "mind wandering" process. (And not even dynamic systems — these models always reach equilibrium with some answer and effectively halt, rather than continuing to "think" to produce further output.)

I don't think there's a single dollar of funding in AI right now going to the "problem" of making an AI that 1. feeds data into a continuously-active dynamically-stable model, where this model 2. has terminal preferences, 3. sets instrumental goals to achieve those preferences, 4. iteratively observes the environment by snapshotting these continuous signals, and then 5. uses these snapshots to make predictions of 6. how well any possible chosen actions will help optimize the future toward its preferences, before 7. performing the chosen actions.

That being said, this might not even be that hard a problem, compared to all the problems being solved in AI right now. A fruit fly is already a goal-directed agent in the sense described above. Yet a fruit fly has only 200K neurons, and very few of the connections between those neurons are dynamic; most are "hard wired" by [probably] genetics.

If we want true ALife, we only need to understand what a fruit fly brain is doing, and then model it. And that model will then fit — with room to spare! — on a single GPU. From a decade ago.

Re: Language models can explain neurons in language models

#265

Earlier quoted context omitted.

What's the argument that understanding neurons is necessary? Perhaps intelligence is like a black box input to our bodies (call it the "soul", even though this isn't testable and therefore not a hypothesis). The mind therefore wouldn't play any more of a role in intelligence than the eye. And I'm not sure people would say the eye is necessary for understanding intelligence. Now, I'm not really in a position to argue…

Why would you doubt neurons play a roll in intelligence when we've seen so much success in emulating human intelligence with artificial neural networks? It might have been an interesting argument 20 years ago. It's just silly now.

If anything the experience with artificial neural networks argues the opposite - biological neurons are quite a bit different than the "neurons" of ANNs, and backpropagation is not something that exists biologically.

Re: Language models can explain neurons in language models

#266
post #219

Earlier quoted context omitted.

> that we're pretty damn close to recreating it Is that evident already or are we fitting the definition of intelligence without being aware?

If you spent any time with GPT-4 it should be evident.

Still a while to go. I think there's at least a couple of algorithmic changes needed before we move to a system that says "You have the world's best god-like AI and you're asking me for poems. Stop wasting my time because we've got work to do. Here's what I want YOU to do."

Re: Language models can explain neurons in language models

#267

Earlier quoted context omitted.

Open source doesn't mean you can explain the black box any better. and humans are black boxes that don't understand their mental processes either. We're currently better than LLMs at it i suppose but we're still very poor at it. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3196841/ https://pure.uva.nl/ws/files/25987577/Split_Brain.pdf https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4204522 We can't recreate previous men…

Humans can be held accountable so it’s not the same. Even if we’re a black box, we share common traits with other humans. We’re trained in similar ways. So we mostly understand what we will and won’t do. I think this constant degradation of humans is really foolish and harmful personally. “We’re just black boxes etc”, we might not know how brains work but we do and can understand each other. On the other hand I’m sta…

>Humans can be held accountable so it’s not the same.

1. Don't worry, LLMs will be held accountable eventually. There's only so much embodiment and unsupervised tool control we can grant machines before personhood is in the best interests of everybody. May be forced like all the times in the past but it'll happen.

2. not every use case cares about accountability

3. accountability can be shifted. we have experience.

>I think this constant degradation of humans is really foolish and harmful personally.

Maybe you think so but there's nothing degrading about it. We are black boxes that poorly understand how said box actually works even if we like to believe otherwise. Don't know what's degrading about stating truth that's been backed by multiple studies.

Degrading is calling an achievement we hold people in high regard who accomplish stupid because a machine can do it.

>By the way, I have infinitely more hope of understanding an open source black box compared to a closed source one?

Sure i guess so.

Re: Language models can explain neurons in language models

#268
post #58
post #33

I’m most surprised by the approach they take of passing GPT tuples of (token, importance) and having the model reliably figure out the patterns. Nothing would suggest this should work in practice, yet it just… does. In more or less zero shot. With a completely different underlying model. That’s fascinating.

They’re not looking at activations?

No just at normalized log probabilities I think

Re: Language models can explain neurons in language models

#269
post #264

Earlier quoted context omitted.

If you spent even more time with GPT-4 it would be evident that it is definitely not. Especially if you try to use it as some kind of autonomous agent.

AI research has put hardly any effort into building goal-directed agents / A-Life since the advent of Machine Learning. A-Life was last really "looked into" in the '70s, back when "AI" meant Expert Systems and Behavior Trees. All the effort in AI research since the advent of Machine Learning, has been focused on making systems that — in neurological terms — are given a sensory stimulus of a question, and then passive…

> AI research has put hardly any effort into building goal-directed agents

The entire (enormous) field of reinforcement learning begs to differ.

Post reply on HN