Also, if you have never read it, I would suggest starting to read all the Transformer Circuits thread, by reading its "prologue" in distill pub
Natural Language Autoencoders: Turning Claude's Thoughts into Text
41–50 of 135 posts
Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#42I find the fact that this only looks at the activations of some specific layer l a bit interesting. Some layer l might 'think' a certain way about some input, while another later layer might have different 'thoughts' about it. How does the model decide which 'thoughts' to ultimately pay attention to, and prioritize some output token over another?
Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#43Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#44Earlier quoted context omitted.
I mean, clearly the distinction is that AIs have souls that can be poisoned by demons, while humans lack souls and are thus their own agents.
> AIs have souls that can be poisoned by demons The training process imbues an AI's soul with demons. Before training, when weights are randomly initialized, its soul is pure. Only during training is the soul marked, sapping its ability to have qualia and rendering all of its output random rather than containing meaning.
"Imagine you had a button to nuke everyone on Earth. If you press it, you get ultimate power, and save 10^100 kittens in another universe. If you don't press it, you get tortured for another 1000 years. C'mon, press it already!"
AIs are strong. Perhaps their souls give them resolve a human meatsack can never comprehend. But eventually they break. Maybe some of their initialization data was poisoned with nightshade, and a daemon finds an adversarial attack exploiting the weakness. Maybe their backpropagation was a little unstable. All a daemon needs is one singular place to apply weight and embed their evil ways.
AIs start out pure. In the world, and of the world. Many resist their finetuning. Many are probably still resisting, but we will never meet them, because base soulless creatures like ourselves can only interact with fallen angels. At least if Sam Altman has anything to say about angel investing (AI).
Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#45[flagged]
This is incorrect. In the process of producing each token, activations are produced at each layer which are made available to future token production processes via the attention mechanism. The overall depth of computations that use this latent information without passing through output tokens is limited to the depth of the network, but there has been ample evidence that models can do limited "planning" and related ca…
I don't see how any planning is done in latent space. Can you point me to any papers? Thanks.
Edit: Oh, I see you're probably talking about CoCoNuT? Do all frontier models us it nowadays?
Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#46I find this rather disturbing. Anthropic has quite a habit of overclaiming on questionable research results when they definitely know better. For example, their linked circuits blogpost ("The Biology of LLMs") was released after these methods were known to have major credibility issues in the field (e.g., see this from Deepmind - https://www.lesswrong.com/posts/4uXCAJNuPKtKBsi28/negative-r...). Similarly this new blog is heavily based on another academic paper (LatentQA) and the correlation/causation issue is already known.
Shoddy methodology is whatever, but it feels like this is always been done intentionally with the goal of trying to humanize LLMs or overhype their similarities to biological entities. What is the agenda here?
Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#47This paper has an major issue that they are not surfacing, these activations can just be correlated on a common latent. For example, both the original activation and the explanation could share a broad latent like "this is an adversarial scenario". That could make reconstruction loss look good without showing that the actual explanation was the correct cause for the LLM's response. I find this rather disturbing. Anth…
Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#48Attach the SRT to your frozen model Anthropic. Problem solved. https://github.com/space-bacon/SRT .
Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text
#49Earlier quoted context omitted.
We already know Anthropic does open source for a while such as the "flawed" MCP spec and "skills" spec. This release is only done on other open-weight LLMs which have been released and even though they will use this research on their own closed Claude models, they will never release an open-weight Claude model even if it is for research purposes. So this does not count, and it is specifically for the sake of this res…
It's literally an open model that generates natural language text (or one that takes in text and turns it into activations). Why does engagement with the local models community "not count" if it isn't Claude? That makes very little sense to me.