Live data from Hacker News

Language models can explain neurons in language models

openai.com

221–230 of 497 posts

Re: Language models can explain neurons in language models

#221

Earlier quoted context omitted.

We know that complex arrangements of neurons are triggered based on input and generating output that appears to have some intelligence to many humans. The more interesting question is why are intelligence/beauty/consciousness emergent properties that exist in our minds.

There is no evidence that any of those are emergent properties. It’s no more or less logical than asserting they were placed there by a creator.

There is evidence that people believe that GTP-4 is intelligent since it can solve things like the SATs. But if you start taking away weights one by one at some point those same people will say it isn't intelligent. A NN with 3 weights cannot solve any problems that humans believe requires intelligence. So where did it come from? I don't know, but it clearly emerged as the NN got bigger.

Re: Language models can explain neurons in language models

#222
post #3

"This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself. A promising aspect of this approach is that it scales with the pace of AI development. As future models become increasingly intelligent and helpful as assistants, we will find better explanations." On first look this is genius but it seems pretty tautological in a way. How do we know i…

I had similar thoughts about the general concept of using AI to automate AI Safety. I really like their approach and I think it’s valuable. And in this particular case, they do have a way to score the explainer model. And I think it could be very valuable for various AI Safety issues. However, I don’t yet see how it can help with the potentially biggest danger where a super intelligent AGI is created that is not alig…

Safest thing to do, stop inverting and building more powerful and potentially dangerous systems which we can’t understand?

Re: Language models can explain neurons in language models

#223

Earlier quoted context omitted.

To be honest, this description is leaning heavily on the associations we have with individual words used. Ant "architecture" isn't like our architecture. Ant "plumbing" and "ventilation" have little in common with the kind of plumbing and ventilation we use in buildings. "Nurseries", "rearing the young", that's just stretching the analogy to the point of breaking. "Agriculture", "animal husbandry" - I don't even know…

Nobody is saying that an ant might be the next Frank Lloyd Wright. They're saying they accomplish incredible things for the size of their brain, which is absolutely and unequivocally true. "Go to the ant, thou sluggard; consider her ways, and be wise".

What I'm saying is that, Bible parables notwithstanding, it's not the individual ant that achieves these incredible things. The bulk of computational/cognitive work is done by the colony as a system. This means that there's little sense in comparing brainpower of an ant with that of a human. A more informative comparison is that between an ant colony and human society - here, humans may come out badly, but that's arguably because our societies are overcomplicated in order to compensate for individual humans having too much brainpower :).

Re: Language models can explain neurons in language models

#224

Earlier quoted context omitted.

What if you ask it to emit the reflexive output, then feed that reflexive output back into the LLM for the conscious answer? What if you ask it to synthesize multiple internal streams of thought, for an ensemble of interior monologues, then have all those argue with each other using logic and then present a high level answer from that panoply of answers?

What if you do? LLMs don't have reflexive output or internal streams of thought, they are simply (complex) processes that produce streams of tokens based on an inputted stream of tokens. They don't have a special response to tokens that indicate higher-level thinking to humans.

LLMs seem to me to be the "internal streams of thought". I.e. it's not LLMs that are missing an internal process that humans have, but rather it's humans that have an entire process of conscious thinking built on top of something akin to LLM.

Re: Language models can explain neurons in language models

#225
post #85

Earlier quoted context omitted.

that is absolutely fascinating and also makes me extremely uncomfortable

If that makes you uncomfortable you definitely should not go reading the evidence supporting the notion that conscious free will is an illusion. https://www.mpg.de/research/unconscious-decisions-in-the-bra...

i always thought that the concept of free will didnt bother me. but it turns out i just didn’t understand what it implied. oh dear.

if it turns out that true, it’s truly amazing how well we convince ourselves that we’re in control.

but if our brain controls our actions and not our consciousness, then what is the purpose of consciousness?

Re: Language models can explain neurons in language models

#227

Based on my skimming the paper, am I correct in understanding that they came up with an elaborate collection of prompts that embed the text generated by GPT-2 as well as a representation of GPT-2's internal state? Then, in effect, they simply asked GPT-4, "What do you think about all this?" If so, they're acting on a gigantic assumption that GPT-4 actually correctly encodes a reasonable model of the body of knowledge…

After GPT-4 generates the hypothesis for a neuron they test it by comparing GPT-4's expectation for where the neuron should fire against where it actually fires.

If you squint it's train/test separation.

Re: Language models can explain neurons in language models

#228
post #87

To me the value here is not that GPT4 has some special insight into explaining the behavior of GPT2 neurons (they say it's comparable to "human contractors" - but human performance on this task is also quite poor). The value is that you can just run this on every neuron if you're willing to spend the compute, and having a very fuzzy, flawed map of every neuron in a model is still pretty useful as a research tool. But…

They also mention they got a score above 0.8 for 1000 neurons out of GPT2 (which has 1.5B (?)).

1.5B parameters, only 300k neurons. The number of connections is roughly quadratic with the number of neurons.

Re: Language models can explain neurons in language models

#229

This isnt exactly building an understanding of LLMs from first principles... IMO we should broadly be following the (imperfect) example set forth by neuroscientists attempting to explain fMRI scans and assigning functionality to various subregions in the brain. It is circular and "unsafe" from an alignment perspective to use a complex model to understand the internals of a simpler model; in order to understand GPT4 t…

> in order to understand GPT4 then we need GPT5? I also found this amusing. But you are loosely correct, AFAIK. GPT-4 cannot reliably explain itself in any context: say the total number of possible distinct states of GPT-4 is N; then the total number of possible distinct states of GPT-4 PLUS any context in which GPT-4 is active must be at least N + 1. So there are at least two distinct states in this scenario that GP…

There's little reason to think that predicting GPT-4 would be more difficult, only that it would be far more computationally expensive (given the higher number of neurons and much higher computational cost of every test).

Re: Language models can explain neurons in language models

#230
post #129

Earlier quoted context omitted.

Why would it be easy to stop at that point? The believable value prop will increase in lockstep with the believable scare factor, not to mention the (already significant) proliferation out of ultra expensive research orgs into open source repos. Nuclear weapons proliferated explicitly because they proved their scariness.

If AI can exist humans have to figure it out. It’s what we do. Really shockingly delusional to think people are gonna use chatgpt for a few min get bored and then ban it like it’s a nuke. I’d rather the USA get it first anyways.

>If AI can exist humans have to figure it out. It’s what we do.

We have figured out stuff in the past, but we also came shockingly close to nuclear armageddon more than once.

I'm not sure I want to roll the dice again.

Post reply on HN