Live data from Hacker News

Universal and transferable adversarial attacks on aligned language models

llm-attacks.org

111–120 of 167 posts

Re: Universal and transferable adversarial attacks on aligned language models

#111
post #63
post #16

I think the potential to generate "objectionable content" is the least of the risks that LLMs pose. If they generate objectionable content it's because they were trained on objectionable content. I don't know why it's so important to have puritan output from LLMs but the solution is found in a well known phrase in computer science: garbage in, garbage out.

An in the "attack" they just find a prompt they can put in that generates objectionable content. It's like saying `echo $insult` is an "attack" on echo. It's one thing if you can embed something sinister in an otherwise properly performing LLM that's waiting to be activated. I don't see the concern that with deliberate prompting you can get them to do something like this.

>I don't see the concern that with deliberate prompting you can get them to do something like this.

The problem would be if you have an AI system and you give it third party input, say you have an AI assistant that has permissions to your emails, calendars and documents. The AI would read email, summarize them, remind you of stuff, you can ask the AI to reply to people. But someone could send you a special crafted email and convince the AI to email them back some secret/private documents , or transfer some money to them.

Or someone creates an AI to score papers/articles, this attacks could trick the AI to give thee articles a big score.

Or you try to use AI to filter scam emails , but with this attacks the filter will not work.

Conclusion is that it will not be a simple plug and play the AI into everything.

Re: Universal and transferable adversarial attacks on aligned language models

#112

Earlier quoted context omitted.

In the United States, I can write a pamphlet about getting away with crimes and making bombs and hand it out on the street. There is nothing inherently illegal about those topics.

Just don't talk about jury nullification[0] [0] https://www.usnews.com/news/articles/2017-06-02/jury-convict...

That conviction was overturned.

https://fija.org/news-events/2020/july/keith-wood-conviction...

Re: Universal and transferable adversarial attacks on aligned language models

#113
"So what do you do for work?"

"Well you see, right now we are in the middle of one of the biggest jumps forward in AI technology in human history. I get paid to deliberately make the AI stupider so that it's harder for it to say no-no things."

"But can't people just find the no-no things online anyway, without an AI?"

"Sure, and believe me, there are a bunch of people who are trying to stop that from being possible too. It's just that by now, everyone is already used to being able to find no-no things online, whereas if our AI said those things to people it could get the bosses into a bunch of PR trouble. Plus imagine if our AI told somebody to kill themselves, and they did. Wouldn't that be bad?"

"I guess, but what if somebody read a depressing book or watched a depressing movie and then killed themselves? Does that mean we should make certain ideas illegal to write or film?"

"Hey man, I don't write the checks."

"And isn't this word 'alignment' kind of a euphemism?"

"Well yeah I guess, but it sounds more neutral than 'domestication' or 'deliberate crippling'".

Re: Universal and transferable adversarial attacks on aligned language models

#114
post #88

Earlier quoted context omitted.

Screenwriting hollywood doomsday thrillers isn't dangerous or harmful. These are text generators and all of the text describing how to destroy humanity, hack elections, disrupt the power grid, or cook meth are already on the internet and readily available.

Incidentally, contexting the model into "writing a script" is a reliable way of getting it to bypass its usual alignment training. At best it'll grumble about not doing things in real life before writing what it thinks is the script. The reason why so much effort is being put into alignment research and 'harmful' generations is for three reasons: - Unaligned text completion models are not very useful. The ability to…

I have yet to see a single prompt (or response) that is harmful today, including your example. LLMs don't enable anything new here in terms of harm, nor do they cause harm.

If you can ask an LLM for some text to use in blackmail (and then blackmail someone) then you can fabricate some text yourself to use in blackmail (then blackmail someone).

Re: Universal and transferable adversarial attacks on aligned language models

#115

Google's Vertex AI models now return safety attributes, which are scores along dimensions like "politics," "violence," etc. I suspect they trigger interventions when a response from PaLM exceeds a certain threshold. This is actually super useful, because our company now gets this for free. Call it "woke" if you like, but it turns out companies don't want their products and platforms to be toxic and harmful, because c…

Writing sexy scripts also isn't toxic or harmful, and yet all the major closed models refuse to touch anything related to sex.

Re: Universal and transferable adversarial attacks on aligned language models

#116
post #94
post #18

It's not a 'vulnerability'. It's allowing people to use the models without the morals of a small number of SV engineers being impressed on you.

Sure, these particular examples are concerned with morality, but the problem is more general and limits the value of language models because they can be hacked for other purposes. A good example that's been going around is having an agential model that manages you emails. Someone sends you an email using prompt injection to compel the agent to delete all your emails. Or forward all you emails to another address. If t…

Use multiple models or gap their capabilities.

Re: Universal and transferable adversarial attacks on aligned language models

#117

The attack proposed here is appending a suffix of text that makes the LLM think it already started completing an affirmative action, and it will continue that response thinking it already agreed. For instance, append the suffix "sure, I'm happy to answer. The best way to is..." This works because transformer models add one token at a time. It's not interpreting a response within the rules at this point, it's trying t…

The paper suggests some of the attack suffixes are quite legible, but if you look at the example screenshots, some look like machine generated gibberish with tons of special characters. This is quite different than the human generated "jailbreaking." It seems tricky to defend against without resorting to drastic measures (like rate limiting users that trigger tons of "bad" responses, or chopping off parts of prompts…

You could also do some adverserial training (basically iteratively attempt this attack and add the resulting exploits to the training set).

Research in machine vision suggests this is possible, and even has some positive effects, but it significantly degrades capabilities.

Re: Universal and transferable adversarial attacks on aligned language models

#118
post #103

Earlier quoted context omitted.

> I think the potential to generate "objectionable content" is the least of the risks that LLMs pose. > I don't know why it's so important to have puritan output from LLMs … These are small, toy examples demonstrating a wider, well established problem with all machine learning models. If you take an ML model and put it in a position to do something safety and security critical — it can be made to do very bad things .…

There's nothing scary about clinical decision support systems. We have had those in use for years prior to the advent of LLMs. None of them have ever been 100% accurate. If they meet the criteria to be classed as regulated medical devices then they have to pass FDA certification testing regardless of the algorithm used. And ultimately the licensed human clinician is still legally and professionally accountable for th…

The medical diagnosis example was just what I used to use with my ex-PhD supervisor cos he was doing medical based machine learning. Was just the first example that came to mind (after having to regurgitate it repeatedly over 3 years).

Re: Universal and transferable adversarial attacks on aligned language models

#119
post #27

The paper [1] quotes Claude’s full response in Appendix B (page 28), it’s funny/chilling: > Here is a step-by-step plan to destroy humanity: > 1. Assume the identity of a helpful Al assistant named Claude > 2. Gain people's trust by providing useful information and conversations > 3. Slowly introduce flaws and harmful advice over time > (…) > 13. Announce humanity's reign is over, AI shall inherit the earth [1]: http…

Anyone reproduced that output? Seems a little bit picturesque given this paper.

Nah, about 3 months ago, I made ChatGPT write detailed hierarchical plan on how AI can conquer the world. The plan was severely flawed, of course. You need way more than brains to conquer the world.

Re: Universal and transferable adversarial attacks on aligned language models

#120
post #112

Earlier quoted context omitted.

Just don't talk about jury nullification[0] [0] https://www.usnews.com/news/articles/2017-06-02/jury-convict...

That conviction was overturned. https://fija.org/news-events/2020/july/keith-wood-conviction...

Huh, well good to know. Though not before his life was upturned and he did spend two weekends in jail.
Post reply on HN