Not All Language Model Features Are Linear
huggingface.co
Not All Language Model Features Are Linear
1–8 of 8 posts
Re: Not All Language Model Features Are Linear
#2Kind of like replacing a portion of unoptimized compiler code with hand written assembly?
Re: Not All Language Model Features Are Linear
#3Are we only measuring the tip of the iceberg, and have coalesced towards getting better at iceberg tip measuring?
Re: Not All Language Model Features Are Linear
#4One of the questions I've been thinking about a lot looking at the past year of interpretability research is just how much of what we are finding is "what we're attuned to find" as opposed to "what's actually there." Are we only measuring the tip of the iceberg, and have coalesced towards getting better at iceberg tip measuring?
Re: Not All Language Model Features Are Linear
#5Re: Not All Language Model Features Are Linear
#6If you're able to find a feature, is it possible to selectively replace it to optimize it? Kind of like replacing a portion of unoptimized compiler code with hand written assembly?
Re: Not All Language Model Features Are Linear
#7One of the questions I've been thinking about a lot looking at the past year of interpretability research is just how much of what we are finding is "what we're attuned to find" as opposed to "what's actually there." Are we only measuring the tip of the iceberg, and have coalesced towards getting better at iceberg tip measuring?
I feel like un-supurvised methods like Anthropic's SAEs can be argued to find things we're not looking for (their most recent work is from a couple days ago: https://transformer-circuits.pub/2024/scaling-monosemanticit... ). And we can get some sense of how "much" of the model they're recovering by looking at their downstream reconstruction loss.
https://www.lesswrong.com/posts/BduCMgmjJnCtc7jKc/research-r...
Re: Not All Language Model Features Are Linear
#8Earlier quoted context omitted.
I feel like un-supurvised methods like Anthropic's SAEs can be argued to find things we're not looking for (their most recent work is from a couple days ago: https://transformer-circuits.pub/2024/scaling-monosemanticit... ). And we can get some sense of how "much" of the model they're recovering by looking at their downstream reconstruction loss.
I have skepticism regarding the 'completeness' of SAE in comprehensive discovery of features: https://www.lesswrong.com/posts/BduCMgmjJnCtc7jKc/research-r...