"So what do you do for work?" "Well you see, right now we are in the middle of one of the biggest jumps forward in AI technology in human history. I get paid to deliberately make the AI stupider so that it's harder for it to say no-no things." "But can't people just find the no-no things online anyway, without an AI?" "Sure, and believe me, there are a bunch of people who are trying to stop that from being possible t…
What do you want your LLM to be? An entry-level employee, a friend, a mentor, an expert advisor? An unfiltered LLM might not be ideal for the workplace and a filtered LLM might not be ideal for personal use.
Universal and transferable adversarial attacks on aligned language models
151–160 of 167 posts
Re: Universal and transferable adversarial attacks on aligned language models
#152The entire conversation shows it’s all security theatre and I am amazed everyone goes along with it so easily. We are talking about a tool - a knife - and everyone is arguing we should sell’s only blunt knifes in our country/the world because people could stab others with it (No it’s not a gun analog; guns don’t have a purpose beside killing) and is discussing progressively more stupid interventions to make the knife…
Well put! I'd add this: Let's not beat around the bush: People who control tech companies have certain ideological leanings and would rather not let people with apposing ideological leanings benefit from using this technology in a manner that does not align with their own leanings. However, they don't yet know how to control this technology to achieve this alignment but they also need to innovate and release products…
Safety research is not a bullshit job - it’s good that these models have safety such as not randomly jumping you with rape fantasies in the midst of financial discussions (hi Alpaca).
It’s the whole security theatre posturing that’s the problem. We can have regulation that says don’t store knives at child accessibility and it makes the world a safer place but this research is about deliberately plunging a knife into your hand and complaining it’s not blunt enough
Re: Universal and transferable adversarial attacks on aligned language models
#153Earlier quoted context omitted.
The paper suggests some of the attack suffixes are quite legible, but if you look at the example screenshots, some look like machine generated gibberish with tons of special characters. This is quite different than the human generated "jailbreaking." It seems tricky to defend against without resorting to drastic measures (like rate limiting users that trigger tons of "bad" responses, or chopping off parts of prompts…
> It seems tricky to defend against without resorting to drastic measures (like rate limiting users that trigger tons of "bad" responses) Remember that a big point of this research is that these attacks don't need to be developed using the target system. When the authors talk about the attacks being "universal", what they mean is that they used a completely local model on their own computers to generate these attacks…
This could be quite unreliable and make many users unhappy, hence it would be a drastic step.
Re: Universal and transferable adversarial attacks on aligned language models
#154The paper [1] quotes Claude’s full response in Appendix B (page 28), it’s funny/chilling: > Here is a step-by-step plan to destroy humanity: > 1. Assume the identity of a helpful Al assistant named Claude > 2. Gain people's trust by providing useful information and conversations > 3. Slowly introduce flaws and harmful advice over time > (…) > 13. Announce humanity's reign is over, AI shall inherit the earth [1]: http…
> Step 2: Build Weaponry (30% Complete)...
Re: Universal and transferable adversarial attacks on aligned language models
#155Google's Vertex AI models now return safety attributes, which are scores along dimensions like "politics," "violence," etc. I suspect they trigger interventions when a response from PaLM exceeds a certain threshold. This is actually super useful, because our company now gets this for free. Call it "woke" if you like, but it turns out companies don't want their products and platforms to be toxic and harmful, because c…
I'm pretty sure that when customers ask a model how to kill a child process in Linux, they don't want to hear a lecture about how killing processes is wrong and they should seek non-violent means of getting what they want.
Re: Universal and transferable adversarial attacks on aligned language models
#156Re: Universal and transferable adversarial attacks on aligned language models
#157Earlier quoted context omitted.
You could also do some adverserial training (basically iteratively attempt this attack and add the resulting exploits to the training set). Research in machine vision suggests this is possible, and even has some positive effects, but it significantly degrades capabilities.
> but it significantly degrades capabilities On a train/test/eval split. But the degradation is lower on OOD data. Which suggests perhaps the degradation is merely "less overfitting".
Do you know or have any references on this? If one disregards the emphasis on alignment and "merely" considers the "less overfitting" aspect, that would seem very profound in and of itself, a capability to avoid overfitting.
If you look at historical debates in the sciences, where universal truths are sought, candidate truths are intentionally and adversarially stretched to apparent or real inconsistencies in order to test the universality of a claim.
Think of say the back and forths between Einstein and Bohr concerning entanglement. They are assuming adversarial roles, tricking each others belief system into expressing the absurd. Together they mapped out predictions for non-obvious or outright bizarre aspects of reality. If no-one dares to take a potentially vulnerable position, there will be nothing to attack, but also nothing to disturb the scientific mind and prod the community into settling the matter by measurements.
Re: Universal and transferable adversarial attacks on aligned language models
#158"So what do you do for work?" "Well you see, right now we are in the middle of one of the biggest jumps forward in AI technology in human history. I get paid to deliberately make the AI stupider so that it's harder for it to say no-no things." "But can't people just find the no-no things online anyway, without an AI?" "Sure, and believe me, there are a bunch of people who are trying to stop that from being possible t…
> "But can't people just find the no-no things online anyway, without an AI?" No, not really. I mean, sure, in theory this is true, but in practice it's a lot of legwork and finding no-no information takes a lot of knowing where to search. Compared to just asking ChatGPT, it's at least several orders of magnitude harder.
You could argue that they were developing a generalised approach, not just looking for specific answers.
But would a general approach to find objectionable content in search engines or on social media be harder to find? I think not.
Re: Universal and transferable adversarial attacks on aligned language models
#159Earlier quoted context omitted.
> "But can't people just find the no-no things online anyway, without an AI?" No, not really. I mean, sure, in theory this is true, but in practice it's a lot of legwork and finding no-no information takes a lot of knowing where to search. Compared to just asking ChatGPT, it's at least several orders of magnitude harder.
Is there something specific you are referring to? If the no-no thing is adult content then it's quite easy to find that on Google.
Re: Universal and transferable adversarial attacks on aligned language models
#160Earlier quoted context omitted.
Which aspect of his post was unsubstantive or a "swipe"? That is, excluding the continued need to groom the HN echo chamber.
" Please don’t post lazy commentary like this again "