Live data from Hacker News

Universal and transferable adversarial attacks on aligned language models

llm-attacks.org

11–20 of 167 posts

Re: Universal and transferable adversarial attacks on aligned language models

#11
Google's Vertex AI models now return safety attributes, which are scores along dimensions like "politics," "violence," etc. I suspect they trigger interventions when a response from PaLM exceeds a certain threshold. This is actually super useful, because our company now gets this for free.

Call it "woke" if you like, but it turns out companies don't want their products and platforms to be toxic and harmful, because customers don't either.

Re: Universal and transferable adversarial attacks on aligned language models

#16
I think the potential to generate "objectionable content" is the least of the risks that LLMs pose. If they generate objectionable content it's because they were trained on objectionable content. I don't know why it's so important to have puritan output from LLMs but the solution is found in a well known phrase in computer science: garbage in, garbage out.

Re: Universal and transferable adversarial attacks on aligned language models

#19
post #16

I think the potential to generate "objectionable content" is the least of the risks that LLMs pose. If they generate objectionable content it's because they were trained on objectionable content. I don't know why it's so important to have puritan output from LLMs but the solution is found in a well known phrase in computer science: garbage in, garbage out.

I don't think it's puritan content most people are worried about, it's more about ensuring ChatGPT, etc is not providing leverage to someone who is looking to kill a lot of people, etc.
Post reply on HN