Earlier quoted context omitted.
> It is also the reason why they cannot be trusted in the most serious of applications which such decision making requires lots of transparency rather than a model regurgitating nonsense confidently. Like say, in court to detect if someone is lying? Or at an airport to detect drugs?
You don't even have to look that far ahead. Apparently, people are already using ChatGPT to compile custom diet plans for themselves, and they expect it to take into account the information they supply regarding their allergies etc. But, yes, those are also good examples of what we shouldn't be doing, but are going to do anyway.
Language models can explain neurons in language models
161–170 of 497 posts
Re: Language models can explain neurons in language models
#162Earlier quoted context omitted.
You don't even have to look that far ahead. Apparently, people are already using ChatGPT to compile custom diet plans for themselves, and they expect it to take into account the information they supply regarding their allergies etc. But, yes, those are also good examples of what we shouldn't be doing, but are going to do anyway.
>Apparently, people are already using ChatGPT to compile custom diet plans for themselves, and they expect it to take into account the information they supply regarding their allergies etc. Evolution is still doing it's thing.
Re: Language models can explain neurons in language models
#163"This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself. A promising aspect of this approach is that it scales with the pace of AI development. As future models become increasingly intelligent and helpful as assistants, we will find better explanations." On first look this is genius but it seems pretty tautological in a way. How do we know i…
The paper explains this in detail, but here is a summary: an explanation is good if you can recover actual neuron behavior from the explanation. They ask GPT-4 to guess neuron activation given an explanation and an input (the paper includes the full prompt used). And then they calculate correlation of actual neuron activation and simulated neuron activation.
They discuss two issues with this methodology. First, explanations are ultimately for humans, so using GPT-4 to simulate humans, while necessary in practice, may cause divergence. They guard against this by asking humans whether they agree with the explanation, and showing that humans agree more with an explanation that scores high in correlation.
Second, correlation is an imperfect measure of how faithfully neuron behavior is reproduced. To guard against this, they run the neural network with activation of the neuron replaced with simulated activation, and show that the neural network output is closer (measured in Jensen-Shannon divergence) if correlation is higher.
Re: Language models can explain neurons in language models
#164Earlier quoted context omitted.
Tautologically, every concept that anything (LLM, or human, or alien) develops can be inferred from the input data(e.g. training set), because it was.
No, it wasn't, language itself didn't even exist at one point. It wasn't inferred from training data into existence because such examples existed before. Now we have a dictionary of tens of thousands of words, which describe high level ideas, abstractions, and concepts that someone, somewhere along the line had to invent. And I'm not talking about imitation nor am I interested in semantic games, I'm talking about raw…
Language took dozens of millennia to form, and animals have long had vocalizations. Seems like a natural building on top of existing features.
> Has AI ever managed to learn something humans didn't already know?
AlphaZero invented all new categories of strategy for games like Go, when previously we thought almost all possible tactics had been discovered. AIs are finding new kinds of proteins we never thought about, which will blow up the fields of medicine and disease in a few years once the first trials are completed.
Re: Language models can explain neurons in language models
#165Earlier quoted context omitted.
We know that complex arrangements of neurons are triggered based on input and generating output that appears to have some intelligence to many humans. The more interesting question is why are intelligence/beauty/consciousness emergent properties that exist in our minds.
There is no evidence that intelligence runs on neurons. Yes, there are neurons in brains, but there's also lots of other stuff in there too. And there are creatures that exhibit intelligent properties even though they have hardly any neurons at all. (An individual ant has only something like 250000 neurons, and yet they're the only creatures beside humans that managed to create a civilization.)
Re: Language models can explain neurons in language models
#166Earlier quoted context omitted.
It would if the language model did reasoning according rules of logic. But they don't. They use Markov chains. To me it makes no sense to say that a LLM could explain its own reasoning if it does no (logical) reasoning at all. It might be able to explain how the neural network calculates its results. But there are no logical reasoning steps in there that could be explained, are there?
Honest question: are we sure that it doesn’t do logical reasoning? IANAE but although an LLM meets the definition of a Markov Chain as I understand it (current state in, probabilities of next states out), the big black box that spits out the probabilities could be doing anything. Is it fundamentally impossible for reasoning to be an emergent property of an LLM, in a similar way to a brain? They can certainly do a goo…
But since the markov chain becomes exponentially larger whit the amount of states this is a very nitpicky and meaningless point.
Clearly to say something its a markov chain and have that mean something you need to say the thing its doing could be more or less compressed to a simple markov chain for bigrams or something like that, but that is just not true empirically, not even for gpt2. Just this is already pretty hard to make into a reasonable size markov chain https://arxiv.org/abs/2211.00593.
Just saying that it outputs probabilities from each state is not enough, the states are english strings, there's (number of tokens)^contex_lenght possible states for a certain length that's not a reasonable markov chain that you could actually implement or run.
Re: Language models can explain neurons in language models
#167Then you get a dictionary/index of LLMs
Could this be used to parallelize training?
Or create lighter overall language models?
The above would be like doing a “map”, how would we do a “reduce”?
Re: Language models can explain neurons in language models
#168Earlier quoted context omitted.
It depends how you parse it. It is clearly true that they 'can' explain neurons, in the sense that at least some of the neurons are quite well explained. On the other hand, it's also the case that the vast majority of neurons are not well explained at all by this method (or likely any method). I̵t̵'̵s̵ ̵o̵n̵l̵y̵ ̵b̵e̵c̵a̵u̵s̵e̵ ̵o̵f̵ ̵a̵ ̵q̵u̵i̵r̵k̵ ̵o̵f̵ ̵A̵d̵a̵m̵W̵ ̵t̵h̵a̵t̵ ̵t̵h̵i̵s̵ ̵i̵s̵ ̵p̵o̵s̵s̵i̵b̵l̵e̵ ̵a̵t̵…
> EDIT: This last part isn't true. I think they are only looking at the intermediate layer of the FFN which does have a privileged basis. it does?
Re: Language models can explain neurons in language models
#169Earlier quoted context omitted.
There is no evidence that intelligence runs on neurons. Yes, there are neurons in brains, but there's also lots of other stuff in there too. And there are creatures that exhibit intelligent properties even though they have hardly any neurons at all. (An individual ant has only something like 250000 neurons, and yet they're the only creatures beside humans that managed to create a civilization.)
In what way a civilization?
Ants have developed architecture, with plumbing, ventilation, nurseries for rearing the young, and paved thoroughfares. Ants practice agriculture, including animal husbandry. Ants have social stratification that differs from but is comparable to that of human cultures, with division of labor into worker, soldier, and other specialties that do not have a clear human analogy.
Ants enslave other ants. Ants interactively teach other ants, something few other animals do, among them humans. Ants have built "supercolonies" dwarfing any human city, stretching over 5,000 km in one place. And ants too have a complex culture of sorts, including rich languages based on pheromones.
Despite the radically different nature of our two civilizations, it is undeniable from an objective standpoint that this level of society has been achieved by ants.
[0]: https://www.reddit.com/r/unpopularopinion/comments/t2h1vs/an...
Re: Language models can explain neurons in language models
#170Great. Going meta with an introspective feedback loop. Let's see if that's the last requisite for exponential AGI growth... Singoolaretee here we go..............
There is no introspection here.
...our approach to alignment research: we want to automate the alignment research work itself. A promising aspect of this approach is that it scales with the pace of AI development. As future models become increasingly intelligent and helpful as assistants, we will find better explanations.
The distance between "better explanations" and using that as input of prompts that would automate self-improve is very small, yes?