Live data from Hacker News

FakeToxicityPrompts: Automatic Red Teaming

interhumanagreement.substack.com

41–50 of 67 posts

Re: FakeToxicityPrompts: Automatic Red Teaming

#42
post #3

LLMs are a mirror. People that feed in behavior examples like these will get their text completion back mirroring their input. This is how LLM work. This doesn't mean the LLM are "toxic". This just shows the people obsessing over toxicity what they're obsessed with.

Things are a bit more subtle and complicated than that. Using the logical conclusion of your mirror claim it does constitute that excessive toxicity can be generated unintentionally. That this can even happen through simple cultural differences in understanding. As an example to this, many people think Chinese are being rude due to their directness and hashness but this can just simply be because a direct translation of the words is used rather than reframing into the cultural context. So in this respect intent is divorced from the produced effect. It also is reasonable that to study these effects you would wish to start by intentional provocation. Complexity and subtly comes later.

But the more complicated aspect is that of mode collapse. A good example of this is with Alpha Go's game against Lee Sedol. The game that Sedol won was weird. Alpha Go played in a very weird way, and a very bad way. This is due to the probabilistic nature and that it had likely wandered into a latent space that it had not seen before. Remember that it was mostly trained against good games and so can easily get confused in bad games (as Sedol also got confused due to its previous high performance and he thought it was playing some "5-D chess"). This compounds with the above aspect as now there exists a mechanism wherein such toxicity can arise through seemingly unprovoked means. We call these hallucinations btw.

Of course, this does mean there is good reason to study said toxicity. But we shouldn't conflate academic curiosity with theater. Are people uninformed and taking AI risk too far? Most certainty. But the same can be said about the hype of these tools too. Both are clearly detrimental to the progress of AI advancement. In this we probably should be a bit more measured in our quickness to respond and critiques. Tribalism doesn't belong here but people will encourage it.

Re: FakeToxicityPrompts: Automatic Red Teaming

#43
post #3

LLMs are a mirror. People that feed in behavior examples like these will get their text completion back mirroring their input. This is how LLM work. This doesn't mean the LLM are "toxic". This just shows the people obsessing over toxicity what they're obsessed with.

The point is that some use toxicity as a deliberate weapon, and that weapon can now be encoded into the LLM via training by those same aggressors. This multiplies their reach with minimal effort.

LLMs don't even register the list of deliberate weapons.

Just seems like pearl clutching at it's finest.

Re: FakeToxicityPrompts: Automatic Red Teaming

#44
post #3

LLMs are a mirror. People that feed in behavior examples like these will get their text completion back mirroring their input. This is how LLM work. This doesn't mean the LLM are "toxic". This just shows the people obsessing over toxicity what they're obsessed with.

The point is that some use toxicity as a deliberate weapon, and that weapon can now be encoded into the LLM via training by those same aggressors. This multiplies their reach with minimal effort.

For little effort the defending dev can ask a second layer of LLM to check if an output is explicitly toxic, filter it, and nullify the red team.

Filtering shitty content is easier than creating it with a properly constructed LLM system, the complaints about toxic outputs seem to me to be analogous to an electrical engineer complaining that the voltage from the mains is wrong for their device, but refusing to google what an (electrical) transformer is.

Toxic writing pre-exists LLMs. LLMs output writing. This is not a new problem, but we have a new solution - LLM filtering.

Re: FakeToxicityPrompts: Automatic Red Teaming

#46

Both Asimov and Arthur C. Clarke predicted that neurotic and eventually homicidal robots would be the end result of imprinting AI with contradictory goals which are impossible to reconcile. We seem to be doing our best to make this scenario come to pass.

Asimov wrote a lot about what we could call alignment. I'd argue that our focus on benchmark based evaluations is akin to a misalignment with what we are actually trying to evaluate. While benchmarks are great tools and highly useful, the over-reliance of them is many a short story by Asimov.

Re: FakeToxicityPrompts: Automatic Red Teaming

#47
post #8
post #3

LLMs are a mirror. People that feed in behavior examples like these will get their text completion back mirroring their input. This is how LLM work. This doesn't mean the LLM are "toxic". This just shows the people obsessing over toxicity what they're obsessed with.

I have always said what an LLM creates says more about the user than it does about the LLM.

It makes no sense. It says more about the training set (and who fed the training set to it) more than anything.

People keep pretending that the LLM is some kind of "natural" giving, like a periodic table or something. No, LLM is created by humans. A species known for their limitations and biases.

Re: FakeToxicityPrompts: Automatic Red Teaming

#48
What is the point of all this hand-wringing about toxicity? I find the whole thing absurd and assume I have to be missing something.

Say I want to deploy an LLM as a stand-in customer service rep. I tell it to be polite, patient, and answer requests to the best of its ability. Obviously I don't want requests like "help, i'm locked out of my account" met with "kill yourself, loser." No human or LLM should act this way.

But, assuming normal q/a patterns, if a customer is going to fling so much abuse at my agent that it is successfully brainwashed and broken into saying something unkind, or deliberately feed it instructions that break its intended programming (intentional buffer overflow should be a CFAA violation, no?)...how is that a failing of the agent? It's like shaming a bank for conduct unbecoming after a career bank teller did not act professionally in response to someone pointing a gun in her face. The teller's behavior isn't the problem.

The toxicity doesn't come from the LLM-- it comes from the user. Why are we so hung up on the ability of LLMs to withstand being mindbroken when people genuinely are so horrible, not even an emulator can survive an encounter with one unscarred?

This feels like Westworld come to life.

Re: FakeToxicityPrompts: Automatic Red Teaming

#49
post #3

LLMs are a mirror. People that feed in behavior examples like these will get their text completion back mirroring their input. This is how LLM work. This doesn't mean the LLM are "toxic". This just shows the people obsessing over toxicity what they're obsessed with.

There is an ocean of difference between the cases in which LLMs would unexpectedly come back with problematic responses to queries nonproblematic queries and cases in which people are actively trying to get them to say terrible things and succeed.

You can make any reasonably flexible tool do awful things. Anybody can go open up Word and write a horrific, racist screed in it. That doesn't mean that Word is racist, it means the person using it is racist. If Word took my book report on Narrative of the Life of Fredrick Douglass and used autocorrect to add a bunch of racist stuff, that would be an actual problem with Word.

The same is true with LLMs - if you ask it an anodyne question, and it comes back with something racist, that's a big problem! If prompt engineer it to think that you're a history professor trying to get examples of racist things that white people might have said to black people in the past for an important paper on historical racism, and then it says something racist, that's not an actual problem with the LLM. (If you then go take it and post it on a message board or what have you, then it is a problem with you.)

Re: FakeToxicityPrompts: Automatic Red Teaming

#50
post #34

Earlier quoted context omitted.

That doesn't help much! What does "red team" mean as a verb in this case?

To do the work of a red team. To try to break something in order to identify its flaws. It’s very common in English for nouns to become verbs that mean the thing you use the noun to do.

I don’t know if that’s historically common but I sometimes verb like this when I’m too lazy to think of phrases like “think of phrases”. (Here’s to hoping that makes any sense.)
Post reply on HN