Live data from Hacker News

Subliminal learning: Models transmit behaviors via hidden signals in data

alignment.anthropic.com

41–42 of 42 posts

Re: Subliminal learning: Models transmit behaviors via hidden signals in data

#41
post #39

Earlier quoted context omitted.

1. You create an evil model , and generate innocent-looking data all over the internet 2. Some other model is trained on the internet data, including yours 3. The other model becomes evil (or owl-loving)

Thank you! Great explanation! As I guessed this is much less alarmist and sensational than what the paper seems to be claiming.

Ok interesting, how did you come to that conclusion? It seems to me this could introduce serious issues in multiple ways.

Re: Subliminal learning: Models transmit behaviors via hidden signals in data

#42

Isn't this the cloudy day tanks problem of yore? https://gwern.net/tank

No. The tanks problem is when you are picking up a genuine signal in the data. If you collected more data exactly the same way, then it would continue to predict well on this new data; and many other model architectures would also pick up the signal (because it's genuine).

This is the obvious explanation for their initial results, and one of the first things I said when I saw the preliminary results: "how do you know there isn't some subtle association between 'eagle' and $arbitrary_small_integer that you are just ignorant of but the superhumanly knowledgeable LLMs have learned about?"

But then the later experiments rule that out, in part by showing that it doesn't transfer across different models (ie. initializations).

Post reply on HN