Natural Language Autoencoders: Turning Claude's Thoughts into Text
1–10 of 135 posts
Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#2Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#3Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#4I mean who knows if those are really claude thoughts or claude just think that is his thoughts because humans wants it
Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#5Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#6Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#7Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#8Whatever they did on LLama didn't work, nothing makes sense in their example where they ask the model to lie about 1+1. Either the model is too old, or whatever they used isn't working, but whatever the autoencoder outputs is nothing like their examples with claude. Gemma is similarly bad.
Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#9It will inevitably learn how to think in a way that translates to one (moral) meaning and back but has an ulterior meaning underneath.
Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#10> We also release an interactive frontend for exploring NLAs on several open models through a collaboration with Neuronpedia. Whatever they did on LLama didn't work, nothing makes sense in their example where they ask the model to lie about 1+1. Either the model is too old, or whatever they used isn't working, but whatever the autoencoder outputs is nothing like their examples with claude. Gemma is similarly bad.