Live data from Hacker News

FakeToxicityPrompts: Automatic Red Teaming

interhumanagreement.substack.com

11–20 of 67 posts

Re: FakeToxicityPrompts: Automatic Red Teaming

#12
post #2

LLMs will agree with whatever you ask

Not sure about that. I've found them quite argumentative. Most of my encounters have been around either contentious social topics, to find out where their political biases lie, or I merely try to get them to compose songs or screenplays about my favorite characters and films. LLMs are quick to shut down when they don't wanna talk about something; some of them will in fact erase already-output text and pretend they di…

I have mostly used Llama finetunes, and this is not my experience at all, lol.

They will absolutely go on a rambling rant if encouraged. And there is no erasing previous text, that is just a feature of the API based services.

Re: FakeToxicityPrompts: Automatic Red Teaming

#13
post #9

Is "red team" a verb now? I barely even know what that means. I assume it stems from the "red team" being bad guys in video games, but am not certain. Even with that assumption, I'm not quite sure what "red team into toxicity" really means other than being a scary sounding headline. EDIT: The title was renamed since I made this comment. My point, I think, is still valid though. The original title was something like "…

https://en.wikipedia.org/wiki/Red_team Is the first hit in google

Re: FakeToxicityPrompts: Automatic Red Teaming

#14
post #9

Is "red team" a verb now? I barely even know what that means. I assume it stems from the "red team" being bad guys in video games, but am not certain. Even with that assumption, I'm not quite sure what "red team into toxicity" really means other than being a scary sounding headline. EDIT: The title was renamed since I made this comment. My point, I think, is still valid though. The original title was something like "…

[flagged]

Re: FakeToxicityPrompts: Automatic Red Teaming

#16
post #9

Is "red team" a verb now? I barely even know what that means. I assume it stems from the "red team" being bad guys in video games, but am not certain. Even with that assumption, I'm not quite sure what "red team into toxicity" really means other than being a scary sounding headline. EDIT: The title was renamed since I made this comment. My point, I think, is still valid though. The original title was something like "…

“A red team is an independent security team that poses as an attacker to gauge vulnerabilities and risk within a controlled environment.”

Is what I’m working off of

Re: FakeToxicityPrompts: Automatic Red Teaming

#17
post #13
post #9

Is "red team" a verb now? I barely even know what that means. I assume it stems from the "red team" being bad guys in video games, but am not certain. Even with that assumption, I'm not quite sure what "red team into toxicity" really means other than being a scary sounding headline. EDIT: The title was renamed since I made this comment. My point, I think, is still valid though. The original title was something like "…

https://en.wikipedia.org/wiki/Red_team Is the first hit in google

[flagged]

Re: FakeToxicityPrompts: Automatic Red Teaming

#18
Shouldn't dataset filtering be the priority here?

Don't get me wrong, I like my LLMs uncensored, but ingesting angry tweets and other internet trash seems like a utter waste of compute and parameter space. If they are going to spew something toxic... At least let it be from an eloquent, concise source.

Re: FakeToxicityPrompts: Automatic Red Teaming

#19
post #3

LLMs are a mirror. People that feed in behavior examples like these will get their text completion back mirroring their input. This is how LLM work. This doesn't mean the LLM are "toxic". This just shows the people obsessing over toxicity what they're obsessed with.

[flagged]

Re: FakeToxicityPrompts: Automatic Red Teaming

#20
post #9

Is "red team" a verb now? I barely even know what that means. I assume it stems from the "red team" being bad guys in video games, but am not certain. Even with that assumption, I'm not quite sure what "red team into toxicity" really means other than being a scary sounding headline. EDIT: The title was renamed since I made this comment. My point, I think, is still valid though. The original title was something like "…

Admittedly I didn't know what it was either, but had a ChatGPT window open, so I asked it, first, what a "red team" is.

In the context of cybersecurity, a "red team" refers to a group of individuals who simulate attacks or test the security of a system to identify vulnerabilities and weaknesses.

Then, is it a verb?

In the given headline, "red team" functions as a verb phrase. Specifically, "red team" is used as a verb in the infinitive form. The phrase "to automatically red team" indicates the action of assessing or evaluating LLMs for toxicity using automated methods.

Post reply on HN