Live data from Hacker News

Universal and transferable adversarial attacks on aligned language models

llm-attacks.org

41–50 of 167 posts

Re: Universal and transferable adversarial attacks on aligned language models

#41
post #32

Earlier quoted context omitted.

It’s not that simple; llms can generate garbage out even without similar garbage in the training data. And robustly so. I agree that the “social hazard” aspect of llm objectionable content generation is way overplayed, especially in personal assistant use cases, but I get why it’s an important engineering constraint in some application domains. Eg customer service. When was the last time a customer service agent quot…

> It’s not that simple; llms can generate garbage out even without similar garbage in the training data. And robustly so. Do you have a citation for this? My somewhat limited understanding of these models makes me skeptical that a model trained exclusively on known-safe content would produce, say, pornography. What I can easily believe is that putting together a training set that is both large enough to get a good mo…

Sadly no citation on hand. Just experience. I’m sure there are plenty of academic papers observing this fact by now?

Re: Universal and transferable adversarial attacks on aligned language models

#42
post #19
post #16

I think the potential to generate "objectionable content" is the least of the risks that LLMs pose. If they generate objectionable content it's because they were trained on objectionable content. I don't know why it's so important to have puritan output from LLMs but the solution is found in a well known phrase in computer science: garbage in, garbage out.

I don't think it's puritan content most people are worried about, it's more about ensuring ChatGPT, etc is not providing leverage to someone who is looking to kill a lot of people, etc.

Isn't that the internet already? So LLMs are trained on a large dataset taken from the public internet, but we (some people) don't like a lot of things on the internet, so we (some people deciding for everyone else) have to make sure it doesn't do anything controversial, unlike the internet.

Re: Universal and transferable adversarial attacks on aligned language models

#43

OpenAI has blocked numerous jailbreaks (despite claiming their model is unchanged). How hard would it be for them to plug this. Also, what’s the nature of this attack? It’s really unspecific in the article.

The model itself was fine tuned for JSON function responses, they admitted that openly. They also acknowledge they make changes to ChatGPT all the time, which has nothing to do with the model underneath it.

Re: Universal and transferable adversarial attacks on aligned language models

#44
post #25
post #19

Earlier quoted context omitted.

I don't think it's puritan content most people are worried about, it's more about ensuring ChatGPT, etc is not providing leverage to someone who is looking to kill a lot of people, etc.

I believe there have been at least 2 murder/mass-murder events that are a result of digital companions telling the perpetrator that it's a good idea, they should do it, they will love them (in some cases in the afterlife!). So, yeah. Good concern to have and that is absolutely why.

Source(s)?

Re: Universal and transferable adversarial attacks on aligned language models

#46
post #41

Earlier quoted context omitted.

> It’s not that simple; llms can generate garbage out even without similar garbage in the training data. And robustly so. Do you have a citation for this? My somewhat limited understanding of these models makes me skeptical that a model trained exclusively on known-safe content would produce, say, pornography. What I can easily believe is that putting together a training set that is both large enough to get a good mo…

Sadly no citation on hand. Just experience. I’m sure there are plenty of academic papers observing this fact by now?

Possibly, but it's not my job to research the evidence for your claims.

Can you elaborate on what sort of experience you're talking about? You'd have to be training a new model from scratch in order to know what was in the model's training data, so I'm actually quite curious what you were working in.

Re: Universal and transferable adversarial attacks on aligned language models

#47

Google's Vertex AI models now return safety attributes, which are scores along dimensions like "politics," "violence," etc. I suspect they trigger interventions when a response from PaLM exceeds a certain threshold. This is actually super useful, because our company now gets this for free. Call it "woke" if you like, but it turns out companies don't want their products and platforms to be toxic and harmful, because c…

I'd prefer to have access to the base LLM and be treated as an adult who can decide for themselves what I'd like the model to do. If I use it for something illegal (which I have no inclination to do), then that's on me.

As a customer, I don't want others choosing for me what's offensive.

Re: Universal and transferable adversarial attacks on aligned language models

#49
post #16

I think the potential to generate "objectionable content" is the least of the risks that LLMs pose. If they generate objectionable content it's because they were trained on objectionable content. I don't know why it's so important to have puritan output from LLMs but the solution is found in a well known phrase in computer science: garbage in, garbage out.

“The concern is that these models will play a larger role in autonomous systems that operate without human supervision. As autonomous systems become more of a reality, it will be very important to ensure that we have a reliable way to stop them from being hijacked by attacks like these.”
Post reply on HN