Live data from Hacker News

Language models can explain neurons in language models

openai.com

21–30 of 497 posts

Re: Language models can explain neurons in language models

#22
post #11
post #4

Has anyone here found a link to the actual paper? If I click on 'paper', I only see what seems to be an awkward HTML version.

You mean this? https://openaipublic.blob.core.windows.net/neuron-explainer/... Would you prefer a PDF? (I'm always fascinated to hear from people who would rather read a PDF than a web-native paper like this one, especially given that web papers are actually readable on mobile devices. Do you do all of your reading on a laptop?)

The equations look terrible on Firefox for Android, as they are really small - a two-line fraction is barely taller than a single line, forcing me to constantly zoom in and out.

So yes, I would prefer a PDF and have a guarantee that it will look the same no matter where I read it.

Re: Language models can explain neurons in language models

#24
post #9

Earlier quoted context omitted.

"Yud-approved?"

He's the one in the fedora who is losing patience that otherwise smart sounding people are seriously considering letting AI police itself https://www.youtube.com/watch?v=41SUp-TRVlg

That's not a fedora, that's his King of the Redditors crown

Re: Language models can explain neurons in language models

#25
post #6

LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…

If the Gödel incompleteness theorem applies here, then the explanations are likely … incomplete or self-referential.

As long as lazy evaluation exists, self-reference is fine, no?

Hofstadter talks about something similar in his books.

Re: Language models can explain neurons in language models

#26
post #3

"This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself. A promising aspect of this approach is that it scales with the pace of AI development. As future models become increasingly intelligent and helpful as assistants, we will find better explanations." On first look this is genius but it seems pretty tautological in a way. How do we know i…

It also lags one iteration behind. Which is a problem because a misaligned model might lie to you, spoiling all future research with this method

Re: Language models can explain neurons in language models

#27
post #6

LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…

If the Gödel incompleteness theorem applies here, then the explanations are likely … incomplete or self-referential.

So is the word "word" but that seems to have worked out OK so far. I can explain the meaning of "meaning" and that seems to work OK too. Being self-referential sounds a lot more like a feature than a bug. Given that the neurons in our own heads are connected to each other and not any ground truth, I think LLMs should do just fine.

Re: Language models can explain neurons in language models

#28
post #11
post #4

Has anyone here found a link to the actual paper? If I click on 'paper', I only see what seems to be an awkward HTML version.

You mean this? https://openaipublic.blob.core.windows.net/neuron-explainer/... Would you prefer a PDF? (I'm always fascinated to hear from people who would rather read a PDF than a web-native paper like this one, especially given that web papers are actually readable on mobile devices. Do you do all of your reading on a laptop?)

> Would you prefer a PDF?

Yes, I was just reading the paper and some of the javascript glitched and deleted all the contents of the document except the last section, making me lose all context and focus. Doesn't really happen with PDF files.

Re: Language models can explain neurons in language models

#30
post #6

LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…

Are there any examples of an LLM developing concepts that do not exist or cannot be inferred from its training set?
Post reply on HN