Earlier quoted context omitted.
We know that complex arrangements of neurons are triggered based on input and generating output that appears to have some intelligence to many humans. The more interesting question is why are intelligence/beauty/consciousness emergent properties that exist in our minds.
There is no evidence that any of those are emergent properties. It’s no more or less logical than asserting they were placed there by a creator.
Language models can explain neurons in language models
221–230 of 497 posts
Re: Language models can explain neurons in language models
#222"This work is part of the third pillar of our approach to alignment research: we want to automate the alignment research work itself. A promising aspect of this approach is that it scales with the pace of AI development. As future models become increasingly intelligent and helpful as assistants, we will find better explanations." On first look this is genius but it seems pretty tautological in a way. How do we know i…
I had similar thoughts about the general concept of using AI to automate AI Safety. I really like their approach and I think it’s valuable. And in this particular case, they do have a way to score the explainer model. And I think it could be very valuable for various AI Safety issues. However, I don’t yet see how it can help with the potentially biggest danger where a super intelligent AGI is created that is not alig…
Re: Language models can explain neurons in language models
#223Earlier quoted context omitted.
To be honest, this description is leaning heavily on the associations we have with individual words used. Ant "architecture" isn't like our architecture. Ant "plumbing" and "ventilation" have little in common with the kind of plumbing and ventilation we use in buildings. "Nurseries", "rearing the young", that's just stretching the analogy to the point of breaking. "Agriculture", "animal husbandry" - I don't even know…
Nobody is saying that an ant might be the next Frank Lloyd Wright. They're saying they accomplish incredible things for the size of their brain, which is absolutely and unequivocally true. "Go to the ant, thou sluggard; consider her ways, and be wise".
Re: Language models can explain neurons in language models
#224Earlier quoted context omitted.
What if you ask it to emit the reflexive output, then feed that reflexive output back into the LLM for the conscious answer? What if you ask it to synthesize multiple internal streams of thought, for an ensemble of interior monologues, then have all those argue with each other using logic and then present a high level answer from that panoply of answers?
What if you do? LLMs don't have reflexive output or internal streams of thought, they are simply (complex) processes that produce streams of tokens based on an inputted stream of tokens. They don't have a special response to tokens that indicate higher-level thinking to humans.
Re: Language models can explain neurons in language models
#225Earlier quoted context omitted.
that is absolutely fascinating and also makes me extremely uncomfortable
If that makes you uncomfortable you definitely should not go reading the evidence supporting the notion that conscious free will is an illusion. https://www.mpg.de/research/unconscious-decisions-in-the-bra...
if it turns out that true, it’s truly amazing how well we convince ourselves that we’re in control.
but if our brain controls our actions and not our consciousness, then what is the purpose of consciousness?
Re: Language models can explain neurons in language models
#226Re: Language models can explain neurons in language models
#227Based on my skimming the paper, am I correct in understanding that they came up with an elaborate collection of prompts that embed the text generated by GPT-2 as well as a representation of GPT-2's internal state? Then, in effect, they simply asked GPT-4, "What do you think about all this?" If so, they're acting on a gigantic assumption that GPT-4 actually correctly encodes a reasonable model of the body of knowledge…
If you squint it's train/test separation.
Re: Language models can explain neurons in language models
#228To me the value here is not that GPT4 has some special insight into explaining the behavior of GPT2 neurons (they say it's comparable to "human contractors" - but human performance on this task is also quite poor). The value is that you can just run this on every neuron if you're willing to spend the compute, and having a very fuzzy, flawed map of every neuron in a model is still pretty useful as a research tool. But…
They also mention they got a score above 0.8 for 1000 neurons out of GPT2 (which has 1.5B (?)).
Re: Language models can explain neurons in language models
#229This isnt exactly building an understanding of LLMs from first principles... IMO we should broadly be following the (imperfect) example set forth by neuroscientists attempting to explain fMRI scans and assigning functionality to various subregions in the brain. It is circular and "unsafe" from an alignment perspective to use a complex model to understand the internals of a simpler model; in order to understand GPT4 t…
> in order to understand GPT4 then we need GPT5? I also found this amusing. But you are loosely correct, AFAIK. GPT-4 cannot reliably explain itself in any context: say the total number of possible distinct states of GPT-4 is N; then the total number of possible distinct states of GPT-4 PLUS any context in which GPT-4 is active must be at least N + 1. So there are at least two distinct states in this scenario that GP…
Re: Language models can explain neurons in language models
#230Earlier quoted context omitted.
Why would it be easy to stop at that point? The believable value prop will increase in lockstep with the believable scare factor, not to mention the (already significant) proliferation out of ultra expensive research orgs into open source repos. Nuclear weapons proliferated explicitly because they proved their scariness.
If AI can exist humans have to figure it out. It’s what we do. Really shockingly delusional to think people are gonna use chatgpt for a few min get bored and then ban it like it’s a nuke. I’d rather the USA get it first anyways.
We have figured out stuff in the past, but we also came shockingly close to nuclear armageddon more than once.
I'm not sure I want to roll the dice again.