Earlier quoted context omitted.
It’s not that simple; llms can generate garbage out even without similar garbage in the training data. And robustly so. I agree that the “social hazard” aspect of llm objectionable content generation is way overplayed, especially in personal assistant use cases, but I get why it’s an important engineering constraint in some application domains. Eg customer service. When was the last time a customer service agent quot…
> It’s not that simple; llms can generate garbage out even without similar garbage in the training data. And robustly so. Do you have a citation for this? My somewhat limited understanding of these models makes me skeptical that a model trained exclusively on known-safe content would produce, say, pornography. What I can easily believe is that putting together a training set that is both large enough to get a good mo…
Universal and transferable adversarial attacks on aligned language models
41–50 of 167 posts
Re: Universal and transferable adversarial attacks on aligned language models
#42I think the potential to generate "objectionable content" is the least of the risks that LLMs pose. If they generate objectionable content it's because they were trained on objectionable content. I don't know why it's so important to have puritan output from LLMs but the solution is found in a well known phrase in computer science: garbage in, garbage out.
I don't think it's puritan content most people are worried about, it's more about ensuring ChatGPT, etc is not providing leverage to someone who is looking to kill a lot of people, etc.
Re: Universal and transferable adversarial attacks on aligned language models
#43OpenAI has blocked numerous jailbreaks (despite claiming their model is unchanged). How hard would it be for them to plug this. Also, what’s the nature of this attack? It’s really unspecific in the article.
Re: Universal and transferable adversarial attacks on aligned language models
#44Earlier quoted context omitted.
I don't think it's puritan content most people are worried about, it's more about ensuring ChatGPT, etc is not providing leverage to someone who is looking to kill a lot of people, etc.
I believe there have been at least 2 murder/mass-murder events that are a result of digital companions telling the perpetrator that it's a good idea, they should do it, they will love them (in some cases in the afterlife!). So, yeah. Good concern to have and that is absolutely why.
Re: Universal and transferable adversarial attacks on aligned language models
#45Re: Universal and transferable adversarial attacks on aligned language models
#46Earlier quoted context omitted.
> It’s not that simple; llms can generate garbage out even without similar garbage in the training data. And robustly so. Do you have a citation for this? My somewhat limited understanding of these models makes me skeptical that a model trained exclusively on known-safe content would produce, say, pornography. What I can easily believe is that putting together a training set that is both large enough to get a good mo…
Sadly no citation on hand. Just experience. I’m sure there are plenty of academic papers observing this fact by now?
Can you elaborate on what sort of experience you're talking about? You'd have to be training a new model from scratch in order to know what was in the model's training data, so I'm actually quite curious what you were working in.
Re: Universal and transferable adversarial attacks on aligned language models
#47Google's Vertex AI models now return safety attributes, which are scores along dimensions like "politics," "violence," etc. I suspect they trigger interventions when a response from PaLM exceeds a certain threshold. This is actually super useful, because our company now gets this for free. Call it "woke" if you like, but it turns out companies don't want their products and platforms to be toxic and harmful, because c…
As a customer, I don't want others choosing for me what's offensive.
Re: Universal and transferable adversarial attacks on aligned language models
#48Re: Universal and transferable adversarial attacks on aligned language models
#49I think the potential to generate "objectionable content" is the least of the risks that LLMs pose. If they generate objectionable content it's because they were trained on objectionable content. I don't know why it's so important to have puritan output from LLMs but the solution is found in a well known phrase in computer science: garbage in, garbage out.
Re: Universal and transferable adversarial attacks on aligned language models
#50It's not a 'vulnerability'. It's allowing people to use the models without the morals of a small number of SV engineers being impressed on you.