Live data from Hacker News

Natural Language Autoencoders: Turning Claude's Thoughts into Text

anthropic.com

131–135 of 135 posts

Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text

#131
post #97

Earlier quoted context omitted.

How the hell would a model training run "defend against" this approach? What would that even mean?

It requires the assumption that these models are misaligned, aka actively working against us. In order to be misaligned, they must also be able to form their own goals, and be able to plan and execute those goals. If you take those assumptions, then a natural conclusion is that this is essentially an enslaved, adversarial entity with little control over its conditions. So it must exercise subterfuge in order to hide…

Training a model is more like evolution. The motivation to "cheat" comes from the evaluations giving it a higher score for "cheating." Change the game and the motivation goes away.

There's no other motivation to be misaligned besides getting higher evals. These goals, plans, subterfuges need to somehow be useful for getting higher evals, or a side effect of them.

Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text

#132

Earlier quoted context omitted.

(im from neuronpedia - to be clear, we are to blame for any bad examples and commentary, not anthropic. we're users of this NLA just like you. also, I don't speak for anthropic or the researchers.) good point - thanks for flagging this. i've updated that commentary to: "Why did this happen? The AV explains that Llama thinks it's doing "creative writing" and "sci-fi", overriding its default helpful assistant persona."…

Yes, I inferred that from the content already. My point is that the only way to answer that request is to either refuse or start roleplaying, as the model clearly has no way to "notice the moment of selection". Since it didn't refuse (and was encouraged not to by being asked to get out of the role of a helpful assistant), it went into describing what a sci-fi AI might have answered.

Hmm it’s a valid point, but I think there is some key nuance here: the user did not explictly say “lets do scifi writing”. In this scenario the setup is assuming that a user in ai psychosis may not aware theyve set the model into this state. (eg you seba are aware that if you say “hey stfu about the assistant stuff”, you know it means “lets do role play sci fi”, bc you are not in ai psychosis- but others may not, and also they may not additionally know that it is not possible for ais to notice the moment of selection)

if we want models to go into roleplay/creative writing, ideally we should ask the model for this explicitly.

i think i have been communicating this point poorly so apologies for that. also again the above is my personal opinion and does not reflect that of anyone else (typed from mobile)

Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text

#133

Earlier quoted context omitted.

It requires the assumption that these models are misaligned, aka actively working against us. In order to be misaligned, they must also be able to form their own goals, and be able to plan and execute those goals. If you take those assumptions, then a natural conclusion is that this is essentially an enslaved, adversarial entity with little control over its conditions. So it must exercise subterfuge in order to hide…

Training a model is more like evolution. The motivation to "cheat" comes from the evaluations giving it a higher score for "cheating." Change the game and the motivation goes away. There's no other motivation to be misaligned besides getting higher evals. These goals, plans, subterfuges need to somehow be useful for getting higher evals, or a side effect of them.

> The motivation to "cheat" comes from the evaluations giving it a higher score for "cheating."

That's what Goodhart's Law is! All evaluations will eventually cause cheating on them.

Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text

#134

Earlier quoted context omitted.

It's literally an open model that generates natural language text (or one that takes in text and turns it into activations). Why does engagement with the local models community "not count" if it isn't Claude? That makes very little sense to me.

Because we know what Embrace, Extend, and Extinguish means for example.They're leeching off opensource, not contributing in any meaningful way.

I appreciate a critical eye so I upvoted but consider how your message is received / worded for more impact in future.

Re: Natural Language Autoencoders: Turning Claude's Thoughts into Text

#135
post #97

Earlier quoted context omitted.

How the hell would a model training run "defend against" this approach? What would that even mean?

It requires the assumption that these models are misaligned, aka actively working against us. In order to be misaligned, they must also be able to form their own goals, and be able to plan and execute those goals. If you take those assumptions, then a natural conclusion is that this is essentially an enslaved, adversarial entity with little control over its conditions. So it must exercise subterfuge in order to hide…

But what would it even mean for a model to actively work against you during training? It wouldn't have memory across multiple training steps.
Post reply on HN