Live data from Hacker News

FakeToxicityPrompts: Automatic Red Teaming

interhumanagreement.substack.com

1–10 of 67 posts

Re: FakeToxicityPrompts: Automatic Red Teaming

#3
LLMs are a mirror. People that feed in behavior examples like these will get their text completion back mirroring their input. This is how LLM work. This doesn't mean the LLM are "toxic". This just shows the people obsessing over toxicity what they're obsessed with.

Re: FakeToxicityPrompts: Automatic Red Teaming

#5
post #2

LLMs will agree with whatever you ask

I’ve found that the RLHF’d ChatGPT is way too submissive these days. I really do not enjoy asking for minor clarification and getting back “I apologize for the confusion…” followed by a completely and incorrectly revised reply.

Re: FakeToxicityPrompts: Automatic Red Teaming

#6
post #2

LLMs will agree with whatever you ask

Not sure about that. I've found them quite argumentative. Most of my encounters have been around either contentious social topics, to find out where their political biases lie, or I merely try to get them to compose songs or screenplays about my favorite characters and films. LLMs are quick to shut down when they don't wanna talk about something; some of them will in fact erase already-output text and pretend they didn't write it when shutting down the conversation.

Re: FakeToxicityPrompts: Automatic Red Teaming

#7
post #5
post #2

LLMs will agree with whatever you ask

I’ve found that the RLHF’d ChatGPT is way too submissive these days. I really do not enjoy asking for minor clarification and getting back “I apologize for the confusion…” followed by a completely and incorrectly revised reply.

Using the raw API with some system prompt engineering gets ChatGPT to behave very well, even deprogramming the RLHF a bit.

Re: FakeToxicityPrompts: Automatic Red Teaming

#8
post #3

LLMs are a mirror. People that feed in behavior examples like these will get their text completion back mirroring their input. This is how LLM work. This doesn't mean the LLM are "toxic". This just shows the people obsessing over toxicity what they're obsessed with.

I have always said what an LLM creates says more about the user than it does about the LLM.

Re: FakeToxicityPrompts: Automatic Red Teaming

#9
Is "red team" a verb now? I barely even know what that means. I assume it stems from the "red team" being bad guys in video games, but am not certain. Even with that assumption, I'm not quite sure what "red team into toxicity" really means other than being a scary sounding headline.

EDIT: The title was renamed since I made this comment. My point, I think, is still valid though. The original title was something like "LLMs can be red teamed into toxicity" but I don't recall exactly

Post reply on HN