I wonder will someone please check the neurons associated to the petertodd and other anomalous glitch tokens ( https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petert... )? I can see the github and I see that for any given neuron you can see associated tokens but I don't see how to do an inverse search.
I think this method will do a poor job at explaining petertodd. These neuron explanations still have to fit within human language for this method to work, and the best you can do in the confines of human language to describe petertodd is to write a long article (just like that link) explaining the many oddities of it. Would be interesting to try, though. I think it's likely that, due to the way glitch tokens happen,…
Language models can explain neurons in language models
391–400 of 497 posts
Re: Language models can explain neurons in language models
#392Earlier quoted context omitted.
Says who? There is actual evidence to support that our brain doesn't "know" what it is doing on a subconscious level? As far as I'm aware it's more that conscious humans don't understand how our brain works. I think the correct statement is "so far the answer is we don't know"
The split brain experiments very very clearly indicate that different parts of the brain can independently conduct behavior and gain knowledge independently of other parts. How or if this generalizes to healthy brains is not super clear, but it does actually provide a good explanatory model for all sorts of self-contradictory behavior (like addiction): the brain has many semi-independent “interests” that are jockeyin…
Neuroimaging indicates high levels of redundancy and variance in spatiotemporal activation.
Studies of cortices and other tissues have already shown that much of the neural tissue of the brain is general purpose.
Why is executive functioning significantly but not exclusively in the tissue of the forebrain, the frontal lobes?
Re: Language models can explain neurons in language models
#393Of note: "... our technique works poorly for larger models, possibly because later layers are harder to explain." And even for GPT-2, which is what they used for the paper: "... the vast majority of our explanations score poorly ..." Which is to say, we still have no clue as to what's going on inside GPT-4 or even GPT-3, which I think is the question many want an answer to. This may be the first step towards that, bu…
Re: Language models can explain neurons in language models
#394LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own. I wonder how many new words we will come up with to describe concepts (or "node-activating clusters of meaning") that the AI finds salient that we don't yet have a singular word for. Or, for that matter, how many of those concepts we will find meaningful at all. What will this teach us about…
"LLMs are quickly going to be able to start explaining their own thought processes better than any human can explain their own." There is no "their" and there is no "thought process" . There is something that produces text that appears to humans like there is something like thought going on (cf the Eliza Effect), but we must be wary of this anthropomorphising language. There is no self reflection, but if you ask an L…
Re: Language models can explain neurons in language models
#395Earlier quoted context omitted.
Do you have any example of how it cannot reason?
It's a neural network. Neural network are not symbolic AI and are not designed to reason
Me: A room contains a blue box and a red box. Alice stands on the red box. Bob also stands on a box. No box has multiple people standing on it. Which box is Bob stood on?
ChatGPT: There are two boxes in the room: a blue box and a red box. Alice is standing on the red box, so Bob must be standing on the blue box. Therefore, Bob is standing on the blue box.
Me: Add a green box to the room. Which box is bob standing on?
ChatGPT: If a green box is added to the room, then there are now three boxes: a blue box, a red box, and a green box. Since Alice is still standing on the red box, we know that Bob must be standing on either the blue box or the green box. We cannot determine which one Bob is standing on without additional information.
Re: Language models can explain neurons in language models
#396Earlier quoted context omitted.
A single human will die permanently, alone, miserable, and unproductive. The comparison is apt.
A single human can survive on their own; there are many historical examples of that. A detached body part, on the other hand, cannot; but it also cannot feel miserable etc. A single ant is more like a body part of the colony in that sense.
But the same is true for an ant.
Re: Language models can explain neurons in language models
#397Re: Language models can explain neurons in language models
#398Earlier quoted context omitted.
Funny that we never quite understood how intelligence worked and yet it appears that we're pretty damn close to recreating it - still without knowing how it works. I wonder how often this happens in the universe...
The battery (Voltaic Pile, 1800) and the telegraph (1830s-1840s) were both invented before the electron was discovered (1897).
Re: Language models can explain neurons in language models
#399Of note: "... our technique works poorly for larger models, possibly because later layers are harder to explain." And even for GPT-2, which is what they used for the paper: "... the vast majority of our explanations score poorly ..." Which is to say, we still have no clue as to what's going on inside GPT-4 or even GPT-3, which I think is the question many want an answer to. This may be the first step towards that, bu…
Funny that we never quite understood how intelligence worked and yet it appears that we're pretty damn close to recreating it - still without knowing how it works. I wonder how often this happens in the universe...
An example: we are not very good at creating flight, the one birds do and humans always regarded as flight, and yet we fly across half the globe in one day.
Going up three meters and landing on a branch is a different matter.
Re: Language models can explain neurons in language models
#400Earlier quoted context omitted.
AI research has put hardly any effort into building goal-directed agents / A-Life since the advent of Machine Learning. A-Life was last really "looked into" in the '70s, back when "AI" meant Expert Systems and Behavior Trees. All the effort in AI research since the advent of Machine Learning, has been focused on making systems that — in neurological terms — are given a sensory stimulus of a question, and then passive…
Brilliant comment—-and back to basics. Yes, and put that compact fruit fly in silico brain into my Roomba please so that it does not get stuck under the bed. This is the kind of embodied AI that should really worry us. Don’t we all suspect deep skunkworks “defense” projects of these types?