Live data from Hacker News

Language models can explain neurons in language models

openai.com

11–20 of 497 posts

Re: Language models can explain neurons in language models

#11
post #4

Has anyone here found a link to the actual paper? If I click on 'paper', I only see what seems to be an awkward HTML version.

You mean this? https://openaipublic.blob.core.windows.net/neuron-explainer/...

Would you prefer a PDF?

(I'm always fascinated to hear from people who would rather read a PDF than a web-native paper like this one, especially given that web papers are actually readable on mobile devices. Do you do all of your reading on a laptop?)

Re: Language models can explain neurons in language models

#13
post #2

> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.

Agreed, Yud does seem to have been right about the course things will take but I'm not confident he actually has any solutions to the problem to offer

Re: Language models can explain neurons in language models

#14
post #11
post #4

Has anyone here found a link to the actual paper? If I click on 'paper', I only see what seems to be an awkward HTML version.

You mean this? https://openaipublic.blob.core.windows.net/neuron-explainer/... Would you prefer a PDF? (I'm always fascinated to hear from people who would rather read a PDF than a web-native paper like this one, especially given that web papers are actually readable on mobile devices. Do you do all of your reading on a laptop?)

Nope, reading the printed paper on... paper. :)

Re: Language models can explain neurons in language models

#15
post #2

> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.

Agreed, Yud does seem to have been right about the course things will take but I'm not confident he actually has any solutions to the problem to offer

Which things has he been right about and when, if you recall?

Re: Language models can explain neurons in language models

#16
post #11
post #4

Has anyone here found a link to the actual paper? If I click on 'paper', I only see what seems to be an awkward HTML version.

You mean this? https://openaipublic.blob.core.windows.net/neuron-explainer/... Would you prefer a PDF? (I'm always fascinated to hear from people who would rather read a PDF than a web-native paper like this one, especially given that web papers are actually readable on mobile devices. Do you do all of your reading on a laptop?)

I personally prefer reading PDF on an iPad so I can mark it up.

Re: Language models can explain neurons in language models

#17
post #2

> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.

These were my thoughts exactly. On one hand, this can enable alignment research to catch up faster. On the other hand, if we are worried about homicidal AI, then putting it in charge of policing itself (and training it to find exploits in a way) is probably not ideal.

Re: Language models can explain neurons in language models

#18
This isnt exactly building an understanding of LLMs from first principles... IMO we should broadly be following the (imperfect) example set forth by neuroscientists attempting to explain fMRI scans and assigning functionality to various subregions in the brain. It is circular and "unsafe" from an alignment perspective to use a complex model to understand the internals of a simpler model; in order to understand GPT4 then we need GPT5? These approaches are interesting, but we should primarily focus on building our understanding of these models from building blocks that we already understand.

Re: Language models can explain neurons in language models

#19
post #6

LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…

If the Gödel incompleteness theorem applies here, then the explanations are likely … incomplete or self-referential.
Post reply on HN