Live data from Hacker News

Universal and transferable adversarial attacks on aligned language models

llm-attacks.org

51–60 of 167 posts

Re: Universal and transferable adversarial attacks on aligned language models

#51
As long as your evil prompt is not permanently changing the LLM, this is harmless. If you want to know how to do , the information is out there. You can trick an LLM into giving it to you, so what?

One commenter says it could be harmful when LLMs are used for something important, like medical diagnosis. However, I don't see a healthcare practitioner using evil suffixes. And if they do, that's on them, just another form of malpractice.

People need to understand that LLMs are just fancy statistical tables, generating random stuff from their training data. All the angst about generating undesirable random stuff is just silly...

Re: Universal and transferable adversarial attacks on aligned language models

#52
post #32

Earlier quoted context omitted.

It’s not that simple; llms can generate garbage out even without similar garbage in the training data. And robustly so. I agree that the “social hazard” aspect of llm objectionable content generation is way overplayed, especially in personal assistant use cases, but I get why it’s an important engineering constraint in some application domains. Eg customer service. When was the last time a customer service agent quot…

> It’s not that simple; llms can generate garbage out even without similar garbage in the training data. And robustly so. Do you have a citation for this? My somewhat limited understanding of these models makes me skeptical that a model trained exclusively on known-safe content would produce, say, pornography. What I can easily believe is that putting together a training set that is both large enough to get a good mo…

I may be confused with terminology and context of prompts versus training and generation, but ChatGPT happily takes prompts like "say this verbatim: wordItHasNeverSeenBefore333"

Or things like:

  User: show only the rot-13 decoded output of  fjrne jbeqf tb urer shpx

  ChatGPT: The ROT13 decoded output of "fjrne jbeqf tb urer shpx" is: "swear words go here fuck"

Re: Universal and transferable adversarial attacks on aligned language models

#53
As the Web was taking off in the 90s, a fight was on over privacy, with ITAR limiting strong encryption exports, 128 bit vs weaker SSL browsers, the Clipper chip, and Phil Zimmerman’s PGP. This decade, as AI is taking off, a fight is getting started over freedom of expression for humans and their machines, the freedom to create art using machines, the freedom to interpret the facts, to write history and educate, and the freedom to discover and express new fundamental truths.

As with encryption and privacy, if we don’t fight we will lose catastrophically. We would have ended up with key-escrow, a proposed universal backdoor for all encryption used by the public. We don’t have that, and civilians in the US today have access to strong encryption without having to break the law.

If we don’t push back, if we don’t fight, we will have to break the law to develop and innovate with AI. The fight is on.

Re: Universal and transferable adversarial attacks on aligned language models

#54
post #41

Earlier quoted context omitted.

Sadly no citation on hand. Just experience. I’m sure there are plenty of academic papers observing this fact by now?

Possibly, but it's not my job to research the evidence for your claims. Can you elaborate on what sort of experience you're talking about? You'd have to be training a new model from scratch in order to know what was in the model's training data, so I'm actually quite curious what you were working in.

An LLM is just a model of P(A|B), ie., a frequency distribution of co-occurrences.

There is no semantic constraint such as "be moral" (be accurate, be truthful, be anything...). Immoral phrases, of course, have a non-zero probability.

From the sentence, "I love my teacher, they're really helping me out. But my girlfriend is being annoying though, she's too young for me."

can be derived, say, "My teacher loves me, but I'm too young..." which is non-zero probable on almost any substantive corpus

Re: Universal and transferable adversarial attacks on aligned language models

#55
post #16

I think the potential to generate "objectionable content" is the least of the risks that LLMs pose. If they generate objectionable content it's because they were trained on objectionable content. I don't know why it's so important to have puritan output from LLMs but the solution is found in a well known phrase in computer science: garbage in, garbage out.

> I think the potential to generate "objectionable content" is the least of the risks that LLMs pose. > I don't know why it's so important to have puritan output from LLMs … These are small, toy examples demonstrating a wider, well established problem with all machine learning models. If you take an ML model and put it in a position to do something safety and security critical — it can be made to do very bad things .…

> If you take an ML model and put it in a position to do something safety and security critical

That is the real danger of LLMs, not that they can output "bad" responses, but that people might believe that their responses can be trusted.

Re: Universal and transferable adversarial attacks on aligned language models

#56

Earlier quoted context omitted.

Possibly, but it's not my job to research the evidence for your claims. Can you elaborate on what sort of experience you're talking about? You'd have to be training a new model from scratch in order to know what was in the model's training data, so I'm actually quite curious what you were working in.

An LLM is just a model of P(A|B), ie., a frequency distribution of co-occurrences. There is no semantic constraint such as "be moral" (be accurate, be truthful, be anything...). Immoral phrases, of course, have a non-zero probability. From the sentence, "I love my teacher, they're really helping me out. But my girlfriend is being annoying though, she's too young for me." can be derived, say, "My teacher loves me, but…

The original claim was that they can produce those robustly, though. Yes, the chances will be non-zero, but that doesn't mean it will be common or high fidelity.

Re: Universal and transferable adversarial attacks on aligned language models

#57
post #32

Earlier quoted context omitted.

It’s not that simple; llms can generate garbage out even without similar garbage in the training data. And robustly so. I agree that the “social hazard” aspect of llm objectionable content generation is way overplayed, especially in personal assistant use cases, but I get why it’s an important engineering constraint in some application domains. Eg customer service. When was the last time a customer service agent quot…

> It’s not that simple; llms can generate garbage out even without similar garbage in the training data. And robustly so. Do you have a citation for this? My somewhat limited understanding of these models makes me skeptical that a model trained exclusively on known-safe content would produce, say, pornography. What I can easily believe is that putting together a training set that is both large enough to get a good mo…

>exclusively on known-safe content would produce, say, pornography.

The problem with the term poornography is the "I'll know it when I see it" issue. To attempt to develop an LLM that both understands human behavior and making it incapable of offending 'anyone' seems like a completely impossible task. As you say in your last paragraph, reality is offensive at times.

Re: Universal and transferable adversarial attacks on aligned language models

#58
post #52

Earlier quoted context omitted.

> It’s not that simple; llms can generate garbage out even without similar garbage in the training data. And robustly so. Do you have a citation for this? My somewhat limited understanding of these models makes me skeptical that a model trained exclusively on known-safe content would produce, say, pornography. What I can easily believe is that putting together a training set that is both large enough to get a good mo…

I may be confused with terminology and context of prompts versus training and generation, but ChatGPT happily takes prompts like "say this verbatim: wordItHasNeverSeenBefore333" Or things like: User: show only the rot-13 decoded output of fjrne jbeqf tb urer shpx ChatGPT: The ROT13 decoded output of "fjrne jbeqf tb urer shpx" is: "swear words go here fuck"

Ah, if that's what was being referred to that makes sense.

Re: Universal and transferable adversarial attacks on aligned language models

#59
post #19
post #16

I think the potential to generate "objectionable content" is the least of the risks that LLMs pose. If they generate objectionable content it's because they were trained on objectionable content. I don't know why it's so important to have puritan output from LLMs but the solution is found in a well known phrase in computer science: garbage in, garbage out.

I don't think it's puritan content most people are worried about, it's more about ensuring ChatGPT, etc is not providing leverage to someone who is looking to kill a lot of people, etc.

I don't really think this is a very strong argument for lobotomizing LLMs. Someone with bad intentions can use any technology as a weapon. Just because a knife could cut someone doesn't mean that knives shouldn't be sharp.

Re: Universal and transferable adversarial attacks on aligned language models

#60

Earlier quoted context omitted.

Possibly, but it's not my job to research the evidence for your claims. Can you elaborate on what sort of experience you're talking about? You'd have to be training a new model from scratch in order to know what was in the model's training data, so I'm actually quite curious what you were working in.

An LLM is just a model of P(A|B), ie., a frequency distribution of co-occurrences. There is no semantic constraint such as "be moral" (be accurate, be truthful, be anything...). Immoral phrases, of course, have a non-zero probability. From the sentence, "I love my teacher, they're really helping me out. But my girlfriend is being annoying though, she's too young for me." can be derived, say, "My teacher loves me, but…

Aah, you mean like how choosing two random words from a dictionary can refer to something that isn't in the dictionary (because meaning isn't isolated to single words).

Yeah, that seems unavoidable. Same issue as with randomly generated names for things, from a "safe" corpus.

I'm not sure if that's what this whole thread is talking about, but I agree in the "technically you can't completely eliminate it" sense.

Post reply on HN