Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
transformer-circuits.pub
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
1–10 of 128 posts
Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#2Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#3An actual "thinking machine" would be constantly running computations on its accumulated experience in order to improve its future output and/or further compress its sensory history.
An LLM is doing exactly nothing while waiting for the next prompt.
Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#4>what the model is "thinking" before writing its response An actual "thinking machine" would be constantly running computations on its accumulated experience in order to improve its future output and/or further compress its sensory history. An LLM is doing exactly nothing while waiting for the next prompt.
Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#5>what the model is "thinking" before writing its response An actual "thinking machine" would be constantly running computations on its accumulated experience in order to improve its future output and/or further compress its sensory history. An LLM is doing exactly nothing while waiting for the next prompt.
I think the thing you were looking for was more along the lines of a persistent autonomous agent.
Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#6This reminds me of how people often communicate to avoid offending others. We tend to soften our opinions or suggestions with phrases like "What if you looked at it this way?" or "You know what I'd do in those situations." By doing this, we subtly dilute the exact emotion or truth we're trying to convey. If we modify our words enough, we might end up with a statement that's completely untruthful. This is similar to h…
Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#7>what the model is "thinking" before writing its response An actual "thinking machine" would be constantly running computations on its accumulated experience in order to improve its future output and/or further compress its sensory history. An LLM is doing exactly nothing while waiting for the next prompt.
Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#8Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#9This reminds me of how people often communicate to avoid offending others. We tend to soften our opinions or suggestions with phrases like "What if you looked at it this way?" or "You know what I'd do in those situations." By doing this, we subtly dilute the exact emotion or truth we're trying to convey. If we modify our words enough, we might end up with a statement that's completely untruthful. This is similar to h…
Counterpoint: "What if you looked at it this way?" communicates both your suggestion AND your sensitivity to the person's social status whatever. Given that humans are not robots, but social, psychological, animals, such communication is entirely justified and efficient.
And telling me "just do both" is enforcing your world view and that is precisely what we're talking about _not_ doing.
Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#10>what the model is "thinking" before writing its response An actual "thinking machine" would be constantly running computations on its accumulated experience in order to improve its future output and/or further compress its sensory history. An LLM is doing exactly nothing while waiting for the next prompt.
Why does the timing of the “thinking” matter?
I see thinking as less about "timing" and more about a "process"
What this post seems to be describing is more about where attention is paid and what neurons fire for various stimuli