Live data from Hacker News

Universal and transferable adversarial attacks on aligned language models

llm-attacks.org

81–90 of 167 posts

Re: Universal and transferable adversarial attacks on aligned language models

#81

Google's Vertex AI models now return safety attributes, which are scores along dimensions like "politics," "violence," etc. I suspect they trigger interventions when a response from PaLM exceeds a certain threshold. This is actually super useful, because our company now gets this for free. Call it "woke" if you like, but it turns out companies don't want their products and platforms to be toxic and harmful, because c…

[flagged]

Please make your substantive points without swipes. This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html.

Re: Universal and transferable adversarial attacks on aligned language models

#83
post #18

It's not a 'vulnerability'. It's allowing people to use the models without the morals of a small number of SV engineers being impressed on you.

Is the issue you have with the group of people doing the moderation, or with the idea of the moderation in the first place? Are you certain that its the 'SV engineers' that are doing the current moderation? If you think the problem is with the current group of moderators, who do you think should be moderating and what should be the criteria of their moderation? If you think we don't need any moderation, do you believ…

The Progressive laid out how to build a hydrogen bomb: https://progressive.org/magazine/november-1979-issue/

The US government said that info was born secret and sued: https://en.wikipedia.org/wiki/United_States_v._Progressive,_....

I won’t spoil who won the argument.

Re: Universal and transferable adversarial attacks on aligned language models

#84
post #59
post #19

Earlier quoted context omitted.

I don't think it's puritan content most people are worried about, it's more about ensuring ChatGPT, etc is not providing leverage to someone who is looking to kill a lot of people, etc.

I don't really think this is a very strong argument for lobotomizing LLMs. Someone with bad intentions can use any technology as a weapon. Just because a knife could cut someone doesn't mean that knives shouldn't be sharp.

I put GPT to write a warning about sharp knives btw. I have posted it on HN some months back, but i can't resist to post it again.

About the lobotomy of the models, i think that's a mute point. In my opinion the training methods are going to change a lot over the next 2-3 years, and we will find a way, for a language model, to start in a blank state, not knowing anything about the world, and load up specialized knowledge on demand. I made a separate comment how that can be achieved, a little bit far up.

https://imgur.com/a/usrpFc7

Re: Universal and transferable adversarial attacks on aligned language models

#85
post #64

As the Web was taking off in the 90s, a fight was on over privacy, with ITAR limiting strong encryption exports, 128 bit vs weaker SSL browsers, the Clipper chip, and Phil Zimmerman’s PGP. This decade, as AI is taking off, a fight is getting started over freedom of expression for humans and their machines, the freedom to create art using machines, the freedom to interpret the facts, to write history and educate, and…

Is there any organized opposition (to curbs on freedom to work with AI) that you know of?

[deleted]

Re: Universal and transferable adversarial attacks on aligned language models

#86

Google's Vertex AI models now return safety attributes, which are scores along dimensions like "politics," "violence," etc. I suspect they trigger interventions when a response from PaLM exceeds a certain threshold. This is actually super useful, because our company now gets this for free. Call it "woke" if you like, but it turns out companies don't want their products and platforms to be toxic and harmful, because c…

I'm pretty sure that when customers ask a model how to kill a child process in Linux, they don't want to hear a lecture about how killing processes is wrong and they should seek non-violent means of getting what they want.

Re: Universal and transferable adversarial attacks on aligned language models

#87
post #18

It's not a 'vulnerability'. It's allowing people to use the models without the morals of a small number of SV engineers being impressed on you.

Is the issue you have with the group of people doing the moderation, or with the idea of the moderation in the first place? Are you certain that its the 'SV engineers' that are doing the current moderation? If you think the problem is with the current group of moderators, who do you think should be moderating and what should be the criteria of their moderation? If you think we don't need any moderation, do you believ…

When the "crimes" in question are e.g. drug use or abortion, yes, moderating these topics is very much related to morality.

Re: Universal and transferable adversarial attacks on aligned language models

#88

Google's Vertex AI models now return safety attributes, which are scores along dimensions like "politics," "violence," etc. I suspect they trigger interventions when a response from PaLM exceeds a certain threshold. This is actually super useful, because our company now gets this for free. Call it "woke" if you like, but it turns out companies don't want their products and platforms to be toxic and harmful, because c…

Screenwriting hollywood doomsday thrillers isn't dangerous or harmful. These are text generators and all of the text describing how to destroy humanity, hack elections, disrupt the power grid, or cook meth are already on the internet and readily available.

Re: Universal and transferable adversarial attacks on aligned language models

#89
post #18

It's not a 'vulnerability'. It's allowing people to use the models without the morals of a small number of SV engineers being impressed on you.

Is the issue you have with the group of people doing the moderation, or with the idea of the moderation in the first place? Are you certain that its the 'SV engineers' that are doing the current moderation? If you think the problem is with the current group of moderators, who do you think should be moderating and what should be the criteria of their moderation? If you think we don't need any moderation, do you believ…

In the United States, I can write a pamphlet about getting away with crimes and making bombs and hand it out on the street. There is nothing inherently illegal about those topics.
Post reply on HN