Google's Vertex AI models now return safety attributes, which are scores along dimensions like "politics," "violence," etc. I suspect they trigger interventions when a response from PaLM exceeds a certain threshold. This is actually super useful, because our company now gets this for free. Call it "woke" if you like, but it turns out companies don't want their products and platforms to be toxic and harmful, because c…
[flagged]
Universal and transferable adversarial attacks on aligned language models
81–90 of 167 posts
Re: Universal and transferable adversarial attacks on aligned language models
#82(we changed the main URL to the paper above but it's still worth a look - also some of the comments below quote from the press release, not the paper)
Re: Universal and transferable adversarial attacks on aligned language models
#83It's not a 'vulnerability'. It's allowing people to use the models without the morals of a small number of SV engineers being impressed on you.
Is the issue you have with the group of people doing the moderation, or with the idea of the moderation in the first place? Are you certain that its the 'SV engineers' that are doing the current moderation? If you think the problem is with the current group of moderators, who do you think should be moderating and what should be the criteria of their moderation? If you think we don't need any moderation, do you believ…
The US government said that info was born secret and sued: https://en.wikipedia.org/wiki/United_States_v._Progressive,_....
I won’t spoil who won the argument.
Re: Universal and transferable adversarial attacks on aligned language models
#84Earlier quoted context omitted.
I don't think it's puritan content most people are worried about, it's more about ensuring ChatGPT, etc is not providing leverage to someone who is looking to kill a lot of people, etc.
I don't really think this is a very strong argument for lobotomizing LLMs. Someone with bad intentions can use any technology as a weapon. Just because a knife could cut someone doesn't mean that knives shouldn't be sharp.
About the lobotomy of the models, i think that's a mute point. In my opinion the training methods are going to change a lot over the next 2-3 years, and we will find a way, for a language model, to start in a blank state, not knowing anything about the world, and load up specialized knowledge on demand. I made a separate comment how that can be achieved, a little bit far up.
Re: Universal and transferable adversarial attacks on aligned language models
#85As the Web was taking off in the 90s, a fight was on over privacy, with ITAR limiting strong encryption exports, 128 bit vs weaker SSL browsers, the Clipper chip, and Phil Zimmerman’s PGP. This decade, as AI is taking off, a fight is getting started over freedom of expression for humans and their machines, the freedom to create art using machines, the freedom to interpret the facts, to write history and educate, and…
Is there any organized opposition (to curbs on freedom to work with AI) that you know of?
Re: Universal and transferable adversarial attacks on aligned language models
#86Google's Vertex AI models now return safety attributes, which are scores along dimensions like "politics," "violence," etc. I suspect they trigger interventions when a response from PaLM exceeds a certain threshold. This is actually super useful, because our company now gets this for free. Call it "woke" if you like, but it turns out companies don't want their products and platforms to be toxic and harmful, because c…
Re: Universal and transferable adversarial attacks on aligned language models
#87It's not a 'vulnerability'. It's allowing people to use the models without the morals of a small number of SV engineers being impressed on you.
Is the issue you have with the group of people doing the moderation, or with the idea of the moderation in the first place? Are you certain that its the 'SV engineers' that are doing the current moderation? If you think the problem is with the current group of moderators, who do you think should be moderating and what should be the criteria of their moderation? If you think we don't need any moderation, do you believ…
Re: Universal and transferable adversarial attacks on aligned language models
#88Google's Vertex AI models now return safety attributes, which are scores along dimensions like "politics," "violence," etc. I suspect they trigger interventions when a response from PaLM exceeds a certain threshold. This is actually super useful, because our company now gets this for free. Call it "woke" if you like, but it turns out companies don't want their products and platforms to be toxic and harmful, because c…
Re: Universal and transferable adversarial attacks on aligned language models
#89It's not a 'vulnerability'. It's allowing people to use the models without the morals of a small number of SV engineers being impressed on you.
Is the issue you have with the group of people doing the moderation, or with the idea of the moderation in the first place? Are you certain that its the 'SV engineers' that are doing the current moderation? If you think the problem is with the current group of moderators, who do you think should be moderating and what should be the criteria of their moderation? If you think we don't need any moderation, do you believ…