It would be interesting to allow users of models to customize inference by tweaking these features, sort of like a semantic equalizer for LLMs. My guess is that this wouldn't work as well as fine-tuning, since that would tweak all the features at once toward your use case, but the equalizer would require zero training data. The prompt itself can trigger the features, so if you say "Try to weave in mentions of San Fra…
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
81–90 of 128 posts
Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#82The article doesn't explain how users can exploit these features in UI or prompt. Does anyone have any insight on how to do so?
Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#83Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#84My observation is, and this may be more philosophical than technical: this process of "decomposing" middle-layer activations with a sparse autoencoder -- is it capturing accurately underlying features in the latent space of the network, or are we drawing order from chaos, imposing monosemanticity where there aren't any? Or to put it another way, were the features always there, learnt by training, or are we doing post-hoc rationalisations -- where the features exist because that's how we defined the autoencoders' dictionaries, and we learn only what we wanted to learn? Are the alien minds of LLMs truly also operating on a similar semantic space as ours, or are we reading tea leaves and seeing what we want to see?
Maybe this distinction doesn't even make sense to begin with; concepts are made by man, if clamping one of these features modifies outputs in a way that is understandable to humans, it doesn't matter if it's capturing some kind of underlying cluster in the latent space of the model. But I do think it's an interesting idea to ponder.
Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#85It would be interesting to allow users of models to customize inference by tweaking these features, sort of like a semantic equalizer for LLMs. My guess is that this wouldn't work as well as fine-tuning, since that would tweak all the features at once toward your use case, but the equalizer would require zero training data. The prompt itself can trigger the features, so if you say "Try to weave in mentions of San Fra…
Over the next year or so I'm sure it will refine enough to be able to be more like a vector multiplier on activation, but simply flipping it on in general is going to create a very 'obsessed' model as stated.
Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#86Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#87I find Anthorpic's work on mech interp fascinating in general. Their initial towards monosemanticity paper was highly surprising, and so is this with the ability to scale to a real production-scale LLM. My observation is, and this may be more philosophical than technical: this process of "decomposing" middle-layer activations with a sparse autoencoder -- is it capturing accurately underlying features in the latent sp…
I'll make a probably bad analogy: does your mindmap place things near each other like my mindmap?
To which I'd say, probably not, mindmaps are very personal, and the more complex we put on ours, the more personal and arbitrary they would be, and the less import the visuals would have
ex. if we have 3 million things on both our mindmaps, it's peering too closely to wonder why you put mcdonalds closer to kids food than restaurants, and you have restaurants in the top left, whereas I put it closer to kids foods, in the top mid left.
Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#88I've really been enjoying their series on mech interp, does anyone have any other good recs?
Was the first research work that clued me into what Anthropic's work today ended up demonstrating.
Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#89I find Anthorpic's work on mech interp fascinating in general. Their initial towards monosemanticity paper was highly surprising, and so is this with the ability to scale to a real production-scale LLM. My observation is, and this may be more philosophical than technical: this process of "decomposing" middle-layer activations with a sparse autoencoder -- is it capturing accurately underlying features in the latent sp…
Re: Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
#90I was pretty upset seeing the superalignment team dissolve at OpenAI, but as is typical for the AI space, the news of one day was quickly eclipsed by the next day.
Anthropic are really killing it right now, and it's very refreshing seeing their commitment to publishing novel findings.
I hope this finally serves as the nail in the coffin on the "it's just fancy autocomplete" and "it doesn't understand what it's saying, bro" rhetoric.