Language models can explain neurons in language models
31–40 of 497 posts
Re: Language models can explain neurons in language models
#32This isnt exactly building an understanding of LLMs from first principles... IMO we should broadly be following the (imperfect) example set forth by neuroscientists attempting to explain fMRI scans and assigning functionality to various subregions in the brain. It is circular and "unsafe" from an alignment perspective to use a complex model to understand the internals of a simpler model; in order to understand GPT4 t…
Re: Language models can explain neurons in language models
#33Nothing would suggest this should work in practice, yet it just… does. In more or less zero shot. With a completely different underlying model. That’s fascinating.
Re: Language models can explain neurons in language models
#34Imagine telling someone in the middle of 2020, that in three years a computer will be able to speak, reason and explain everything as if it was a human, absolutely incredible!
Re: Language models can explain neurons in language models
#35LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…
If the Gödel incompleteness theorem applies here, then the explanations are likely … incomplete or self-referential.
Re: Language models can explain neurons in language models
#36Has anyone here found a link to the actual paper? If I click on 'paper', I only see what seems to be an awkward HTML version.
You mean this? https://openaipublic.blob.core.windows.net/neuron-explainer/... Would you prefer a PDF? (I'm always fascinated to hear from people who would rather read a PDF than a web-native paper like this one, especially given that web papers are actually readable on mobile devices. Do you do all of your reading on a laptop?)
Re: Language models can explain neurons in language models
#37> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.
Agreed, Yud does seem to have been right about the course things will take but I'm not confident he actually has any solutions to the problem to offer
Re: Language models can explain neurons in language models
#38Re: Language models can explain neurons in language models
#39LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…
If the Gödel incompleteness theorem applies here, then the explanations are likely … incomplete or self-referential.
Of course, if you were trying to use GPT-4 to explain GPT-4 then I think the Gödel incompleteness theorem would be more relevant, and even then I'm not so sure.
Re: Language models can explain neurons in language models
#40LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…
Are there any examples of an LLM developing concepts that do not exist or cannot be inferred from its training set?