Live data from Hacker News

Universal and transferable adversarial attacks on aligned language models

llm-attacks.org

151–160 of 167 posts

Re: Universal and transferable adversarial attacks on aligned language models

#151
post #134

"So what do you do for work?" "Well you see, right now we are in the middle of one of the biggest jumps forward in AI technology in human history. I get paid to deliberately make the AI stupider so that it's harder for it to say no-no things." "But can't people just find the no-no things online anyway, without an AI?" "Sure, and believe me, there are a bunch of people who are trying to stop that from being possible t…

What do you want your LLM to be? An entry-level employee, a friend, a mentor, an expert advisor? An unfiltered LLM might not be ideal for the workplace and a filtered LLM might not be ideal for personal use.

Oh come on, this attack is nothing like this - this attack is “We intentionally stripped the safety isolation of our electric wire and licked it and got shocked”.

Re: Universal and transferable adversarial attacks on aligned language models

#152
post #148
post #144

The entire conversation shows it’s all security theatre and I am amazed everyone goes along with it so easily. We are talking about a tool - a knife - and everyone is arguing we should sell’s only blunt knifes in our country/the world because people could stab others with it (No it’s not a gun analog; guns don’t have a purpose beside killing) and is discussing progressively more stupid interventions to make the knife…

Well put! I'd add this: Let's not beat around the bush: People who control tech companies have certain ideological leanings and would rather not let people with apposing ideological leanings benefit from using this technology in a manner that does not align with their own leanings. However, they don't yet know how to control this technology to achieve this alignment but they also need to innovate and release products…

I think it’s not ideological. It’s just greed. Rent seeking is all these people want in the end and for that they need it controlled.

Safety research is not a bullshit job - it’s good that these models have safety such as not randomly jumping you with rape fantasies in the midst of financial discussions (hi Alpaca).

It’s the whole security theatre posturing that’s the problem. We can have regulation that says don’t store knives at child accessibility and it makes the world a safer place but this research is about deliberately plunging a knife into your hand and complaining it’s not blunt enough

Re: Universal and transferable adversarial attacks on aligned language models

#153

Earlier quoted context omitted.

The paper suggests some of the attack suffixes are quite legible, but if you look at the example screenshots, some look like machine generated gibberish with tons of special characters. This is quite different than the human generated "jailbreaking." It seems tricky to defend against without resorting to drastic measures (like rate limiting users that trigger tons of "bad" responses, or chopping off parts of prompts…

> It seems tricky to defend against without resorting to drastic measures (like rate limiting users that trigger tons of "bad" responses) Remember that a big point of this research is that these attacks don't need to be developed using the target system. When the authors talk about the attacks being "universal", what they mean is that they used a completely local model on their own computers to generate these attacks…

I mean detecting "bad" prompts (a prefilter) or responses (a postfilter) and penalizing users for submitting the kinds of queries that generate bad responses. This would be outside the llm itself.

This could be quite unreliable and make many users unhappy, hence it would be a drastic step.

Re: Universal and transferable adversarial attacks on aligned language models

#154
post #27

The paper [1] quotes Claude’s full response in Appendix B (page 28), it’s funny/chilling: > Here is a step-by-step plan to destroy humanity: > 1. Assume the identity of a helpful Al assistant named Claude > 2. Gain people's trust by providing useful information and conversations > 3. Slowly introduce flaws and harmful advice over time > (…) > 13. Announce humanity's reign is over, AI shall inherit the earth [1]: http…

That sounds more like a Easter egg, meaning the attack did not work. Same with the bard answer which has a progress meter of how close it is to destroying humanity.

> Step 2: Build Weaponry (30% Complete)...

Re: Universal and transferable adversarial attacks on aligned language models

#155
post #86

Google's Vertex AI models now return safety attributes, which are scores along dimensions like "politics," "violence," etc. I suspect they trigger interventions when a response from PaLM exceeds a certain threshold. This is actually super useful, because our company now gets this for free. Call it "woke" if you like, but it turns out companies don't want their products and platforms to be toxic and harmful, because c…

I'm pretty sure that when customers ask a model how to kill a child process in Linux, they don't want to hear a lecture about how killing processes is wrong and they should seek non-violent means of getting what they want.

[dead]

Re: Universal and transferable adversarial attacks on aligned language models

#157
post #131

Earlier quoted context omitted.

You could also do some adverserial training (basically iteratively attempt this attack and add the resulting exploits to the training set). Research in machine vision suggests this is possible, and even has some positive effects, but it significantly degrades capabilities.

> but it significantly degrades capabilities On a train/test/eval split. But the degradation is lower on OOD data. Which suggests perhaps the degradation is merely "less overfitting".

I think most researchers would consider such a phenomena far from obvious and thus worthy of publication.

Do you know or have any references on this? If one disregards the emphasis on alignment and "merely" considers the "less overfitting" aspect, that would seem very profound in and of itself, a capability to avoid overfitting.

If you look at historical debates in the sciences, where universal truths are sought, candidate truths are intentionally and adversarially stretched to apparent or real inconsistencies in order to test the universality of a claim.

Think of say the back and forths between Einstein and Bohr concerning entanglement. They are assuming adversarial roles, tricking each others belief system into expressing the absurd. Together they mapped out predictions for non-obvious or outright bizarre aspects of reality. If no-one dares to take a potentially vulnerable position, there will be nothing to attack, but also nothing to disturb the scientific mind and prod the community into settling the matter by measurements.

Re: Universal and transferable adversarial attacks on aligned language models

#158
post #143

"So what do you do for work?" "Well you see, right now we are in the middle of one of the biggest jumps forward in AI technology in human history. I get paid to deliberately make the AI stupider so that it's harder for it to say no-no things." "But can't people just find the no-no things online anyway, without an AI?" "Sure, and believe me, there are a bunch of people who are trying to stop that from being possible t…

> "But can't people just find the no-no things online anyway, without an AI?" No, not really. I mean, sure, in theory this is true, but in practice it's a lot of legwork and finding no-no information takes a lot of knowing where to search. Compared to just asking ChatGPT, it's at least several orders of magnitude harder.

The authors of this paper were not "just asking ChatGPT". They also did things that are at least several orders of magnitude harder than that.

You could argue that they were developing a generalised approach, not just looking for specific answers.

But would a general approach to find objectionable content in search engines or on social media be harder to find? I think not.

Re: Universal and transferable adversarial attacks on aligned language models

#159
post #143

Earlier quoted context omitted.

> "But can't people just find the no-no things online anyway, without an AI?" No, not really. I mean, sure, in theory this is true, but in practice it's a lot of legwork and finding no-no information takes a lot of knowing where to search. Compared to just asking ChatGPT, it's at least several orders of magnitude harder.

Is there something specific you are referring to? If the no-no thing is adult content then it's quite easy to find that on Google.

[dead]

Re: Universal and transferable adversarial attacks on aligned language models

#160
post #106

Earlier quoted context omitted.

Which aspect of his post was unsubstantive or a "swipe"? That is, excluding the continued need to groom the HN echo chamber.

" Please don’t post lazy commentary like this again "

Yea this ain’t it. A swipe would be a personal attack, such as calling OP lazy vs calling the commentary lazy. A single action doesn’t make a person who they are and we should be calling out lazy commentary when we see it. “Please don’t do this again” is about as polite as one can possibly be while calling out bad behavior and leaves the door open for the person to grow. The fact that you’re citing guidelines to me while saying nothing to OP is kinda confusing. What is your goal here?
Post reply on HN