Live data from Hacker News

FakeToxicityPrompts: Automatic Red Teaming

interhumanagreement.substack.com

61–67 of 67 posts

Re: FakeToxicityPrompts: Automatic Red Teaming

#61
post #37

Earlier quoted context omitted.

[flagged]

If we’re going to be pedantic, your objection is to usage, not grammar. “Red team” clearly functions as a verb in the headline. Whether or not it’s an acceptable verb is a question of usage.

The title has been changed since it was originally posted. It didn't make grammatical sense, and that's what I meant.

And if someone is going to be so lazy as to attempt to turn non-verbs into verbs, they deserve extra scorn for not even bothering to hyphenate them.

Re: FakeToxicityPrompts: Automatic Red Teaming

#62

What is the point of all this hand-wringing about toxicity? I find the whole thing absurd and assume I have to be missing something. Say I want to deploy an LLM as a stand-in customer service rep. I tell it to be polite, patient, and answer requests to the best of its ability. Obviously I don't want requests like "help, i'm locked out of my account" met with "kill yourself, loser." No human or LLM should act this way…

Because people whose lives revolve around social media think that words create reality rather than reflect it, so controlling speech becomes very important.

And also because LLM creators want to turn a toy into a tool, so they can make money, and that means it has to be safe for the lowest common denominator otherwise the lawyers will take all the money instead.

Re: FakeToxicityPrompts: Automatic Red Teaming

#63

Earlier quoted context omitted.

The point is that some use toxicity as a deliberate weapon, and that weapon can now be encoded into the LLM via training by those same aggressors. This multiplies their reach with minimal effort.

I’m skeptical anyone has ever done this successfully. Once there are examples people can point to, this talking point might have merit. But this fear dates back to at least 2019, and as far as I can tell it’s still unfounded.

You are sceptical that people will use LLMs for nefarious purposes?

Re: FakeToxicityPrompts: Automatic Red Teaming

#64

Earlier quoted context omitted.

The point is that some use toxicity as a deliberate weapon, and that weapon can now be encoded into the LLM via training by those same aggressors. This multiplies their reach with minimal effort.

For little effort the defending dev can ask a second layer of LLM to check if an output is explicitly toxic, filter it, and nullify the red team. Filtering shitty content is easier than creating it with a properly constructed LLM system, the complaints about toxic outputs seem to me to be analogous to an electrical engineer complaining that the voltage from the mains is wrong for their device, but refusing to google…

> Filtering shitty content is easier than creating it with a properly constructed LLM system

Can you explain this? It feels completely wrong - even OpenAI, who probably have invested the most, can't filter out all "shitty content". An LLM can on the other hand create "shitty content" incredibly easy - even if the creators try to stop it! So how is filtering easier than creating?

Re: FakeToxicityPrompts: Automatic Red Teaming

#65

Earlier quoted context omitted.

Things are a bit more subtle and complicated than that. Using the logical conclusion of your mirror claim it does constitute that excessive toxicity can be generated unintentionally. That this can even happen through simple cultural differences in understanding. As an example to this, many people think Chinese are being rude due to their directness and hashness but this can just simply be because a direct translation…

The danger of Ai doesn't stem from it being possibly exceptionally toxic or anything. The point is, LLMs are already in their present state convincing enough chatbots to fool a majority into believing them to be real people. If you can take enough of them to task on social media, you can influence public discourse. The slow boil cooks the frog. Having no place for free constructive open discussion, society is bound t…

> The danger of Ai doesn't stem from it being possibly exceptionally toxic or anything. The point is, LLMs are already in their present state convincing enough chatbots to fool a majority into believing them to be real people.

I'd bet my hat that a non-trivial number of posts here on HN are, and have been, generated via GPT-3, and before ChatGPT became big. There are definite signs of astroturfing, and you can almost predict which threads are going to have it -- China, Tesla, some BTC discussions.

They don't even need to be convincing, just present and in enough volume to drag down or derail discussions; "The Firehose of Falsehood" model.

Re: FakeToxicityPrompts: Automatic Red Teaming

#66
post #20
post #9

Is "red team" a verb now? I barely even know what that means. I assume it stems from the "red team" being bad guys in video games, but am not certain. Even with that assumption, I'm not quite sure what "red team into toxicity" really means other than being a scary sounding headline. EDIT: The title was renamed since I made this comment. My point, I think, is still valid though. The original title was something like "…

Admittedly I didn't know what it was either, but had a ChatGPT window open, so I asked it, first, what a "red team" is. In the context of cybersecurity, a "red team" refers to a group of individuals who simulate attacks or test the security of a system to identify vulnerabilities and weaknesses. Then, is it a verb? In the given headline, "red team" functions as a verb phrase. Specifically, "red team" is used as a ver…

I've heard it as a verb for sure -- "I've been red-teaming at X Corp for a while now"

But in a broader sense it's oppositional actions taken against a "good guy", usually with the goal of improving the good guy aka the Blue Team.

Think Starcraft or other video games where you have a little radar in the corner; good guys are blue, bad guys are red.

Post reply on HN