Live data from Hacker News

Language models can explain neurons in language models

openai.com

31–40 of 497 posts

Re: Language models can explain neurons in language models

#32

This isnt exactly building an understanding of LLMs from first principles... IMO we should broadly be following the (imperfect) example set forth by neuroscientists attempting to explain fMRI scans and assigning functionality to various subregions in the brain. It is circular and "unsafe" from an alignment perspective to use a complex model to understand the internals of a simpler model; in order to understand GPT4 t…

I don't follow. Neuroscience imaging tools like fMRI are only used because it is impossible to measure the activations of each neuron in a brain in real time (unlike an artificial neural network). This research paper's attempt to understand the role of individual neurons or neuron clusters within a complete network gets much closer to "first principles" than fMRI.

Re: Language models can explain neurons in language models

#33
I’m most surprised by the approach they take of passing GPT tuples of (token, importance) and having the model reliably figure out the patterns.

Nothing would suggest this should work in practice, yet it just… does. In more or less zero shot. With a completely different underlying model. That’s fascinating.

Re: Language models can explain neurons in language models

#34
For people overwhelmed by all the AI science speak, just spend a few minutes with bing or phind and it will explain everything surprisingly well.

Imagine telling someone in the middle of 2020, that in three years a computer will be able to speak, reason and explain everything as if it was a human, absolutely incredible!

Re: Language models can explain neurons in language models

#35
post #6

LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…

If the Gödel incompleteness theorem applies here, then the explanations are likely … incomplete or self-referential.

The Goedel Incompleteness Theorem has no straightforward application to this question.

Re: Language models can explain neurons in language models

#36
post #11
post #4

Has anyone here found a link to the actual paper? If I click on 'paper', I only see what seems to be an awkward HTML version.

You mean this? https://openaipublic.blob.core.windows.net/neuron-explainer/... Would you prefer a PDF? (I'm always fascinated to hear from people who would rather read a PDF than a web-native paper like this one, especially given that web papers are actually readable on mobile devices. Do you do all of your reading on a laptop?)

If you want to draw on it, PDF is usually the best

Re: Language models can explain neurons in language models

#37
post #2

> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.

Agreed, Yud does seem to have been right about the course things will take but I'm not confident he actually has any solutions to the problem to offer

His solution is a global regulatory regime to ban new large training runs. The tools required to accomplish this are, IMO, out of the question but I will give Yud credit for being honest about them while others who share his viewpoint try to hide the ball.

Re: Language models can explain neurons in language models

#39
post #6

LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…

If the Gödel incompleteness theorem applies here, then the explanations are likely … incomplete or self-referential.

That's probably one of the reasons why you'd use GPT-4 to explain GPT-2.

Of course, if you were trying to use GPT-4 to explain GPT-4 then I think the Gödel incompleteness theorem would be more relevant, and even then I'm not so sure.

Re: Language models can explain neurons in language models

#40
post #6

LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…

Are there any examples of an LLM developing concepts that do not exist or cannot be inferred from its training set?

The training sets are so poorly curated we will never know...
Post reply on HN