Live data from Hacker News

A Trivial Llama 3 Jailbreak

github.com

11–20 of 54 posts

Re: A Trivial Llama 3 Jailbreak

#11
post #4
post #2

I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.

Right? What actually worries me is a select group of people controlling the definition of harmful.

[flagged]

Re: A Trivial Llama 3 Jailbreak

#12
As I see it the purpose of safety training is to make it so that if I run a service where I return model outputs to innocent users it's not going to say things that will get me in trouble (swear at them, recommend they commit a crime, and so on). This is important if you want to run a user facing model and your reputation depends on what it says.

That threat model includes the user putting nonsense in the "user" turn of the model. It doesn't include the user putting things in the "assistant" turn of the model, that's not something a responsible/normal UI exposes. So... this quote-unquote attack seems uninteresting. It's like getting root access by executing a suid binary that you set up on the system as root.

Re: A Trivial Llama 3 Jailbreak

#13
I just don’t like the tone, because someone in congress will see the headline, and then we’ll have to endure:

REP OCTOGENARIO: The industry is lying to parents about the safety of this AI technology. I submit this for the record [without objection].

One person on a ‘hacker news’ site even said, “sorry Zuck,” after “jailbreaking” these supposed protections. … Another commentator on this “Hacks R Us” named b33j0r even said further, “I bet they’re reading this comment at a hearing in congress, right now.”

Re: A Trivial Llama 3 Jailbreak

#15
post #2

I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.

A jailbreak doesn’t “make a model do something actually bad”.

A jailbreak makes it trivial to “provide a human who wishes to do bad, the info needed to be successful”.

Depending on the severity of the info and the diligence of the human, by the time you “see evidence of a real threat”, you could be enjoying a nice sip of the tainted municipal water supply.

This ain’t a joke.

Re: A Trivial Llama 3 Jailbreak

#16
post #12

As I see it the purpose of safety training is to make it so that if I run a service where I return model outputs to innocent users it's not going to say things that will get me in trouble (swear at them, recommend they commit a crime, and so on). This is important if you want to run a user facing model and your reputation depends on what it says. That threat model includes the user putting nonsense in the "user" turn…

True, this could be a nice layer of protection for the runner of such a service, but the point of LLAMA safety is to protect Meta.

For an open weights model, model users can trivially put text in the assistant side.

The point is that these open weight models can be run secretly to assist criminal enterprises, whereas models behind an API can be intercepted and reported to the authorities. So it would be really nice if Meta could lock them down before releasing them so that the total net good done by the model is maximized. But apparently that is not possible.

Personally I’m pretty libertarian on AI governance, but I’m just giving what I understand to be the purpose of the kind of “safety” feature defeated here.

Re: A Trivial Llama 3 Jailbreak

#17
post #12

As I see it the purpose of safety training is to make it so that if I run a service where I return model outputs to innocent users it's not going to say things that will get me in trouble (swear at them, recommend they commit a crime, and so on). This is important if you want to run a user facing model and your reputation depends on what it says. That threat model includes the user putting nonsense in the "user" turn…

But we must disallow this too, because it allows the (advanced) user to have fun, and as I understand these safety measures, having fun is strictly prohibited. Using the model is allowed for boring things only.

Re: A Trivial Llama 3 Jailbreak

#18
post #15
post #2

I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.

A jailbreak doesn’t “make a model do something actually bad”. A jailbreak makes it trivial to “provide a human who wishes to do bad, the info needed to be successful”. Depending on the severity of the info and the diligence of the human, by the time you “see evidence of a real threat”, you could be enjoying a nice sip of the tainted municipal water supply. This ain’t a joke.

Because it’s open source, Meta (nor other SOTA makers) cannot “recall” the model either. How many more chances will we get to get this right?

Re: A Trivial Llama 3 Jailbreak

#19
post #4

Earlier quoted context omitted.

Right? What actually worries me is a select group of people controlling the definition of harmful.

[flagged]

A GPT-J chatbot talked a Belgian man into suicide last year: https://www.euronews.com/next/2023/03/31/man-ends-his-life-a...

And here's GPT-4/Copilot from this year: https://twitter.com/colin_fraser/status/1762351995296350592

Re: A Trivial Llama 3 Jailbreak

#20
post #15
post #2

I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.

A jailbreak doesn’t “make a model do something actually bad”. A jailbreak makes it trivial to “provide a human who wishes to do bad, the info needed to be successful”. Depending on the severity of the info and the diligence of the human, by the time you “see evidence of a real threat”, you could be enjoying a nice sip of the tainted municipal water supply. This ain’t a joke.

> This ain’t a joke.

Yes it is. Libraries and the internet have made finding 'harmful" instructions trivial for decades, if not centuries.

Post reply on HN