Live data from Hacker News

Language models can explain neurons in language models

openai.com

281–290 of 497 posts

Re: Language models can explain neurons in language models

#281
post #3

"This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself. A promising aspect of this approach is that it scales with the pace of AI development. As future models become increasingly intelligent and helpful as assistants, we will find better explanations." On first look this is genius but it seems pretty tautological in a way. How do we know i…

> How do we know if the explainer is good? The paper explains this in detail, but here is a summary: an explanation is good if you can recover actual neuron behavior from the explanation. They ask GPT-4 to guess neuron activation given an explanation and an input (the paper includes the full prompt used). And then they calculate correlation of actual neuron activation and simulated neuron activation. They discuss two…

> The paper explains this in detail, but here is a summary: an explanation is good if you can recover actual neuron behavior from the explanation.

To be clear, this is only neuron activation strength for text inputs. We aren't doing any mechanistic modeling of whether our explanation of what the neuron does predicts any role the neuron might play within the internals of the network, despite most neurons likely having a role that can only be succinctly summarized in relation to the rest of the network.

It seems very easy to end up with explanations that correlate well with a neuron, but do not actually meaningfully explain what the neuron is doing.

Re: Language models can explain neurons in language models

#282

Earlier quoted context omitted.

There is no evidence that intelligence runs on neurons. Yes, there are neurons in brains, but there's also lots of other stuff in there too. And there are creatures that exhibit intelligent properties even though they have hardly any neurons at all. (An individual ant has only something like 250000 neurons, and yet they're the only creatures beside humans that managed to create a civilization.)

What else would intelligence run on?

Every cell in our body, and every bacterium living in a body (e.g. gut flora), contribute to our intelligence. It looks plausible (to me) that there's one "top cell" among them that represents the "person", others just contributing via layered signals, but whether this "top cell" is a neuron or another kind of cell is unknown.

Re: Language models can explain neurons in language models

#283
post #249

Earlier quoted context omitted.

If you spent even more time with GPT-4 it would be evident that it is definitely not. Especially if you try to use it as some kind of autonomous agent.

What have you tried to do with it?

Use it to analyze the California & US Code, the California & Federal Codes of Regulation, and bills currently in the California legislation & Congress. It's far from useless but far more useful for creative writing than any kind of understanding or instruction following when it comes to complex topics.

Even performing a map-reduce over large documents to summarize or analyze them for a specific audience is largely beyond it. A 32K context size is a pittance when it comes to a single Title in the USC or CFR, which average into the millions of tokens each.

Re: Language models can explain neurons in language models

#284
post #219

Earlier quoted context omitted.

If you spent any time with GPT-4 it should be evident.

If you spent even more time with GPT-4 it would be evident that it is definitely not. Especially if you try to use it as some kind of autonomous agent.

I think we'll soon be able to train models that answer any reasonable question. By that measure, computers are intelligent, and getting smarter by the day. But I don't think that is the bar we care about. In the context of intelligence, I believe we care about self-directed thought, or agency. And a computer program needs to keep running to achieve that because it needs to interact with the world.

Re: Language models can explain neurons in language models

#285
post #219

Earlier quoted context omitted.

If you spent any time with GPT-4 it should be evident.

Its vast limitations in anything reasoning-based are indeed evident.

GPT-4 is better at reasoning than 90% of humans. At least. I won't be surprised if GPT-5 is better than 100% of humans. I'm saying this in complete seriousness.

Re: Language models can explain neurons in language models

#286

Earlier quoted context omitted.

What if you do? LLMs don't have reflexive output or internal streams of thought, they are simply (complex) processes that produce streams of tokens based on an inputted stream of tokens. They don't have a special response to tokens that indicate higher-level thinking to humans.

LLMs seem to me to be the "internal streams of thought". I.e. it's not LLMs that are missing an internal process that humans have, but rather it's humans that have an entire process of conscious thinking built on top of something akin to LLM.

Well put, and I agree. My belief is that if a typical person was drugged or otherwise induced to just blurt out their unfiltered thoughts out loud as it crossed their brain, the level of incohesion and false confidence on display would look a lot like an LLM hallucinating.

Re: Language models can explain neurons in language models

#287
post #89

Of note: "... our technique works poorly for larger models, possibly because later layers are harder to explain." And even for GPT-2, which is what they used for the paper: "... the vast majority of our explanations score poorly ..." Which is to say, we still have no clue as to what's going on inside GPT-4 or even GPT-3, which I think is the question many want an answer to. This may be the first step towards that, bu…

Funny that we never quite understood how intelligence worked and yet it appears that we're pretty damn close to recreating it - still without knowing how it works. I wonder how often this happens in the universe...

Yep, we don't know all constituents of buttermilk, nor how bread stales (there's too much going on inside). But it doesn't prevent us to judge their usefulness.

Re: Language models can explain neurons in language models

#288

Seems like OpenAI is grasping at straws trying to make GPT "go meta". Reminds me of this Sam Altman quote from 2019: "We have made a soft promise to investors that once we build this sort-of generally intelligent system, basically we will ask it to figure out a way to generate an investment return." https://youtu.be/TzcJlKg2Rc0?t=1886

I have a similar feeling, they’ve potentially built the most amazing but commercially useless thing in history. I don’t mean it’s not useful entirely, but I mean. It’s not useful in that it’s not deterministic enough to be trustworthy, it’s dangerous and really hard to scale therefore it’s more of an academic project than something that will make Altman as famous as Sergey Brin. I personally take people like Hinton s…

Time will tell. Anecdotally, I know several professional who find ChatGPT3.5 & 4 to be valuable and willing to pay for access. I certainly save more than $20 per month for my work by using ChatGPT to accelerate my day to day activities.

Re: Language models can explain neurons in language models

#289

Earlier quoted context omitted.

If you spent even more time with GPT-4 it would be evident that it is definitely not. Especially if you try to use it as some kind of autonomous agent.

If you spent even more time with GPT-4 it would be evident that it definitely is. Especialy if you try to use it as some kind of autonomous agent. (Notice how baseless comments can sway either way)

Engaging with this is probably a mistake, but remember the burden of proof is on the claimant. What examples do you have of ChatGPT for example, learning in a basic classroom setting, or navigating an escape room, or being inspired to create its own spontaneous art, or founding a startup, or…

Re: Language models can explain neurons in language models

#290
post #285

Earlier quoted context omitted.

Its vast limitations in anything reasoning-based are indeed evident.

GPT-4 is better at reasoning than 90% of humans. At least. I won't be surprised if GPT-5 is better than 100% of humans. I'm saying this in complete seriousness.

Do you put yourself in the 10% or the 90%? I’m asking in complete seriousness.
Post reply on HN