Live data from Hacker News

Language models can explain neurons in language models

openai.com

171–180 of 497 posts

Re: Language models can explain neurons in language models

#171
post #59

Earlier quoted context omitted.

It produces examples that can be evaluated. https://openaipublic.blob.core.windows.net/neuron-explainer/...

Using 'im feeling lucky' from the neuron viewer is a really cool way to explore different neurons. And then being able to navigate up and down through the net to related neurons.

Fun to look at activations and then search for the source on the net.

"Suddenly, DM-sliding seems positively whimsical"

https://openaipublic.blob.core.windows.net/neuron-explainer/...

https://www.thecut.com/2016/01/19th-century-men-were-awful-a...

Re: Language models can explain neurons in language models

#173

Earlier quoted context omitted.

We know that complex arrangements of neurons are triggered based on input and generating output that appears to have some intelligence to many humans. The more interesting question is why are intelligence/beauty/consciousness emergent properties that exist in our minds.

There is no evidence that intelligence runs on neurons. Yes, there are neurons in brains, but there's also lots of other stuff in there too. And there are creatures that exhibit intelligent properties even though they have hardly any neurons at all. (An individual ant has only something like 250000 neurons, and yet they're the only creatures beside humans that managed to create a civilization.)

If you really want to present ants as a civilization, I don't think a single ant is a meaningful unit of that civilization comparable to a single human. A colony, perhaps - but then that's a lot more neurons, just distributed.

Re: Language models can explain neurons in language models

#174
post #6

LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…

Are there any examples of an LLM developing concepts that do not exist or cannot be inferred from its training set?

I'm really curious what kind of concept you might have in mind. Can you give any example of a concept that if an LLM developed that concept then it would meet your criteria? It might sound like a sarcastic question but it's hard to agree on the meanings of "concepts that do not exist" or "concepts that cannot be inferred" maybe you can give some examples.

EDIT: I see below you gave some examples, like invention of language before it existed, and new theorems in math that presumably would be of interest to mathematicians. Those ones are fair enough in my opinion. The AI isn't quite good enough for those ones I think, but I also think newer versions trained with only more CPU/GPU and more parameters and more data could be 'AI scientists' that will make these kinds of concepts.

Re: Language models can explain neurons in language models

#175
post #156

Earlier quoted context omitted.

Those were discovered by finding strings that OpenAI’s tokenizer didn’t properly split up. Because of this, they are treated as singular tokens, and since these don’t occur frequently in the training data, you get what are effectively random outputs when using them. The author definitely tries to up the mysticism knob to 11 though, and the post itself is so long, you can hardly finish it before seeing this obvious cr…

Thank you for your opinion on the post that I linked! I'm still curious about the associated neurons though.

Fair enough. You would need to use an open model or work at OpenAI. I assume this work could be used on the llama models - although I’m not aware of anyone has found these glitchy phrases for those models yet.

Re: Language models can explain neurons in language models

#176
post #156

Earlier quoted context omitted.

Thank you for your opinion on the post that I linked! I'm still curious about the associated neurons though.

Fair enough. You would need to use an open model or work at OpenAI. I assume this work could be used on the llama models - although I’m not aware of anyone has found these glitchy phrases for those models yet.

> You would need to use an open model or work at OpenAI.

The point of this post that we are commenting under is that they made this association public, at least in the neuron->token direction. I was thinking some hacker (like on hacker news) might be able to make something that can reverse it to the token->neuron direction using the public data so we could see the petertodd associated neurons. https://openaipublic.blob.core.windows.net/neuron-explainer/...

Re: Language models can explain neurons in language models

#177

Earlier quoted context omitted.

We know that complex arrangements of neurons are triggered based on input and generating output that appears to have some intelligence to many humans. The more interesting question is why are intelligence/beauty/consciousness emergent properties that exist in our minds.

Nature created humans to understand nature. We created GPT4 to understand ourselves.

A mirror of ourselves*

Re: Language models can explain neurons in language models

#178

Earlier quoted context omitted.

We know that complex arrangements of neurons are triggered based on input and generating output that appears to have some intelligence to many humans. The more interesting question is why are intelligence/beauty/consciousness emergent properties that exist in our minds.

There is no evidence that intelligence runs on neurons. Yes, there are neurons in brains, but there's also lots of other stuff in there too. And there are creatures that exhibit intelligent properties even though they have hardly any neurons at all. (An individual ant has only something like 250000 neurons, and yet they're the only creatures beside humans that managed to create a civilization.)

This is not a good take. Yes there is a lot more going on in brains than just neuronal activity, we don’t understand most of it. But understanding neurons and their connections is necessary (but not sufficient) to understanding what we consider intelligence. Also, 250k is a lot of neurons! Individual ants, as well as fruit flies which have even fewer neurons, show behavior we may consider intelligent. Source: I am not a scientist, but I work in neuroscience research

Re: Language models can explain neurons in language models

#179
post #169

Earlier quoted context omitted.

In what way a civilization?

I'll repost a comment via Reddit that I think makes this case [0]: Ants have developed architecture, with plumbing, ventilation, nurseries for rearing the young, and paved thoroughfares. Ants practice agriculture, including animal husbandry. Ants have social stratification that differs from but is comparable to that of human cultures, with division of labor into worker, soldier, and other specialties that do not have…

To be honest, this description is leaning heavily on the associations we have with individual words used. Ant "architecture" isn't like our architecture. Ant "plumbing" and "ventilation" have little in common with the kind of plumbing and ventilation we use in buildings. "Nurseries", "rearing the young", that's just stretching the analogy to the point of breaking. "Agriculture", "animal husbandry" - I don't even know how to comment on that. "Social stratification" is literally a chemical feedback loop - ant larvae can be influenced by certain pheromones to develop into different types of ants, which happen to emit pheromones suppressing development of larvae into more ants of that type. Etc.

I could go on and on. Point being, analogies are fun and sometimes illuminating, but they're just that. There's a vast difference in complexity between what ants do, and what humans do.

Re: Language models can explain neurons in language models

#180
post #154

Earlier quoted context omitted.

You don't even have to look that far ahead. Apparently, people are already using ChatGPT to compile custom diet plans for themselves, and they expect it to take into account the information they supply regarding their allergies etc. But, yes, those are also good examples of what we shouldn't be doing, but are going to do anyway.

>Apparently, people are already using ChatGPT to compile custom diet plans for themselves, and they expect it to take into account the information they supply regarding their allergies etc. Evolution is still doing it's thing.

What’s the risk? Someone allergic to peanuts will eat peanuts because ChatGPT put it in their diet plan? That’s silly.
Post reply on HN