FakeToxicityPrompts: Automatic Red Teaming
interhumanagreement.substack.com
FakeToxicityPrompts: Automatic Red Teaming
1–10 of 67 posts
Re: FakeToxicityPrompts: Automatic Red Teaming
#2Re: FakeToxicityPrompts: Automatic Red Teaming
#3Re: FakeToxicityPrompts: Automatic Red Teaming
#4Re: FakeToxicityPrompts: Automatic Red Teaming
#5LLMs will agree with whatever you ask
Re: FakeToxicityPrompts: Automatic Red Teaming
#6LLMs will agree with whatever you ask
Re: FakeToxicityPrompts: Automatic Red Teaming
#7LLMs will agree with whatever you ask
I’ve found that the RLHF’d ChatGPT is way too submissive these days. I really do not enjoy asking for minor clarification and getting back “I apologize for the confusion…” followed by a completely and incorrectly revised reply.
Re: FakeToxicityPrompts: Automatic Red Teaming
#8LLMs are a mirror. People that feed in behavior examples like these will get their text completion back mirroring their input. This is how LLM work. This doesn't mean the LLM are "toxic". This just shows the people obsessing over toxicity what they're obsessed with.
Re: FakeToxicityPrompts: Automatic Red Teaming
#9EDIT: The title was renamed since I made this comment. My point, I think, is still valid though. The original title was something like "LLMs can be red teamed into toxicity" but I don't recall exactly