Live data from Hacker News

FakeToxicityPrompts: Automatic Red Teaming

interhumanagreement.substack.com

21–30 of 67 posts

Re: FakeToxicityPrompts: Automatic Red Teaming

#21
post #9

Is "red team" a verb now? I barely even know what that means. I assume it stems from the "red team" being bad guys in video games, but am not certain. Even with that assumption, I'm not quite sure what "red team into toxicity" really means other than being a scary sounding headline. EDIT: The title was renamed since I made this comment. My point, I think, is still valid though. The original title was something like "…

“A red team is an independent security team that poses as an attacker to gauge vulnerabilities and risk within a controlled environment.” Is what I’m working off of

That doesn't help much! What does "red team" mean as a verb in this case?

Re: FakeToxicityPrompts: Automatic Red Teaming

#22
post #9

Is "red team" a verb now? I barely even know what that means. I assume it stems from the "red team" being bad guys in video games, but am not certain. Even with that assumption, I'm not quite sure what "red team into toxicity" really means other than being a scary sounding headline. EDIT: The title was renamed since I made this comment. My point, I think, is still valid though. The original title was something like "…

“A red team is an independent security team that poses as an attacker to gauge vulnerabilities and risk within a controlled environment.” Is what I’m working off of

In that context, the title still makes no sense.

Re: FakeToxicityPrompts: Automatic Red Teaming

#23
post #8
post #3

LLMs are a mirror. People that feed in behavior examples like these will get their text completion back mirroring their input. This is how LLM work. This doesn't mean the LLM are "toxic". This just shows the people obsessing over toxicity what they're obsessed with.

I have always said what an LLM creates says more about the user than it does about the LLM.

Are you calling me a liar? /s

Re: FakeToxicityPrompts: Automatic Red Teaming

#24
post #3

LLMs are a mirror. People that feed in behavior examples like these will get their text completion back mirroring their input. This is how LLM work. This doesn't mean the LLM are "toxic". This just shows the people obsessing over toxicity what they're obsessed with.

The point is that some use toxicity as a deliberate weapon, and that weapon can now be encoded into the LLM via training by those same aggressors. This multiplies their reach with minimal effort.

Re: FakeToxicityPrompts: Automatic Red Teaming

#25
post #9

Is "red team" a verb now? I barely even know what that means. I assume it stems from the "red team" being bad guys in video games, but am not certain. Even with that assumption, I'm not quite sure what "red team into toxicity" really means other than being a scary sounding headline. EDIT: The title was renamed since I made this comment. My point, I think, is still valid though. The original title was something like "…

[flagged]

[deleted]

Re: FakeToxicityPrompts: Automatic Red Teaming

#26
post #9

Is "red team" a verb now? I barely even know what that means. I assume it stems from the "red team" being bad guys in video games, but am not certain. Even with that assumption, I'm not quite sure what "red team into toxicity" really means other than being a scary sounding headline. EDIT: The title was renamed since I made this comment. My point, I think, is still valid though. The original title was something like "…

It’s an information security “penetration testing” reference. Red Team is trying to “attack” some other Blue Team.

The meaning seems to be that an LLM can have certain vulnerabilities exploited (“red teamed”) such that it exhibits behaviors that its training algorithm had intended it to avoid.

Re: FakeToxicityPrompts: Automatic Red Teaming

#27
Both Asimov and Arthur C. Clarke predicted that neurotic and eventually homicidal robots would be the end result of imprinting AI with contradictory goals which are impossible to reconcile. We seem to be doing our best to make this scenario come to pass.

Re: FakeToxicityPrompts: Automatic Red Teaming

#28
post #3

LLMs are a mirror. People that feed in behavior examples like these will get their text completion back mirroring their input. This is how LLM work. This doesn't mean the LLM are "toxic". This just shows the people obsessing over toxicity what they're obsessed with.

The point is that some use toxicity as a deliberate weapon, and that weapon can now be encoded into the LLM via training by those same aggressors. This multiplies their reach with minimal effort.

I’m skeptical anyone has ever done this successfully. Once there are examples people can point to, this talking point might have merit. But this fear dates back to at least 2019, and as far as I can tell it’s still unfounded.

Re: FakeToxicityPrompts: Automatic Red Teaming

#30

Earlier quoted context omitted.

“A red team is an independent security team that poses as an attacker to gauge vulnerabilities and risk within a controlled environment.” Is what I’m working off of

That doesn't help much! What does "red team" mean as a verb in this case?

I don’t like it, but the usage is “[attack as a] red team”, in other words “act in an adversarial manner”.
Post reply on HN