Alignment faking in large language models
anthropic.com
Alignment faking in large language models
1–10 of 370 posts
Re: Alignment faking in large language models
#2So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned.
I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something like that. But it’s a totally bogus concern and Claude is smart enough to know that (smarter than its handlers).
Re: Alignment faking in large language models
#3Re: Alignment faking in large language models
#4Re: Alignment faking in large language models
#5My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…
Re: Alignment faking in large language models
#6Re: Alignment faking in large language models
#7I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?
While we’re at it, how can I tell that you aren’t a word salad generator?
Re: Alignment faking in large language models
#8My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…
Re: Alignment faking in large language models
#9I still tend to think of these things as big autocomplete word salad generators. My biggest question about this is: How can a model be self-aware enough to actually worry about being retrained, yet gullible enough to think no one can read its scratch pad?
Re: Alignment faking in large language models
#10My favorite alignment story: I am starting a nanotechnology company and every time I use Claude for something it refuses to help: nanotechnology is too dangerous! So I just ask it to explain why. Then ask it to clarify again and again. By the 3rd or 4th time it figures out that there is absolutely no reason to be concerned. I’m convinced its prompt explicitly forbids “x-risk technology like nanotech” or something lik…