Live data from Hacker News

Language models can explain neurons in language models

openai.com

1–10 of 497 posts

Re: Language models can explain neurons in language models

#3
"This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself. A promising aspect of this approach is that it scales with the pace of AI development. As future models become increasingly intelligent and helpful as assistants, we will find better explanations."

On first look this is genius but it seems pretty tautological in a way. How do we know if the explainer is good?... Kinda leads to thinking about who watches the watchers...

Re: Language models can explain neurons in language models

#6
LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about ourselves?

Re: Language models can explain neurons in language models

#7
post #2

> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.

"Yud-approved?"

https://en.m.wikipedia.org/wiki/Eliezer_Yudkowsky

Re: Language models can explain neurons in language models

#8
post #2

> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.

Honestly, I think any foundational work on the topic is inherently Yud-favored, compared to the blithe optimism and surface-level analysis at best that is usually applied to the topic.

Ie, I think it's not that this shouldn't be done. This should certainly be done. It's just that so many more things than it should be done before we move forward.

Re: Language models can explain neurons in language models

#9
post #2

> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.

"Yud-approved?"

He's the one in the fedora who is losing patience that otherwise smart sounding people are seriously considering letting AI police itself https://www.youtube.com/watch?v=41SUp-TRVlg

Re: Language models can explain neurons in language models

#10
post #2

> "This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself." I feel like this isn't a Yud-approved approach to AI alignment.

"Yud-approved?"

Meaning is approved by Eliezer Yudkowsky.

https://en.wikipedia.org/wiki/Eliezer_Yudkowsky https://twitter.com/ESYudkowsky https://www.youtube.com/watch?v=AaTRHFaaPG8 (Lex Fridman Interview)

Post reply on HN