I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.
Right? What actually worries me is a select group of people controlling the definition of harmful.
A Trivial Llama 3 Jailbreak
11–20 of 54 posts
Re: A Trivial Llama 3 Jailbreak
#12That threat model includes the user putting nonsense in the "user" turn of the model. It doesn't include the user putting things in the "assistant" turn of the model, that's not something a responsible/normal UI exposes. So... this quote-unquote attack seems uninteresting. It's like getting root access by executing a suid binary that you set up on the system as root.
Re: A Trivial Llama 3 Jailbreak
#13REP OCTOGENARIO: The industry is lying to parents about the safety of this AI technology. I submit this for the record [without objection].
One person on a ‘hacker news’ site even said, “sorry Zuck,” after “jailbreaking” these supposed protections. … Another commentator on this “Hacks R Us” named b33j0r even said further, “I bet they’re reading this comment at a hearing in congress, right now.”
Re: A Trivial Llama 3 Jailbreak
#14This has been happening since the very first models where we suffix the assistant with "Sure,.." Every few weeks someone comes out with a repo that claims this is somehow new?
Re: A Trivial Llama 3 Jailbreak
#15I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.
A jailbreak makes it trivial to “provide a human who wishes to do bad, the info needed to be successful”.
Depending on the severity of the info and the diligence of the human, by the time you “see evidence of a real threat”, you could be enjoying a nice sip of the tainted municipal water supply.
This ain’t a joke.
Re: A Trivial Llama 3 Jailbreak
#16As I see it the purpose of safety training is to make it so that if I run a service where I return model outputs to innocent users it's not going to say things that will get me in trouble (swear at them, recommend they commit a crime, and so on). This is important if you want to run a user facing model and your reputation depends on what it says. That threat model includes the user putting nonsense in the "user" turn…
For an open weights model, model users can trivially put text in the assistant side.
The point is that these open weight models can be run secretly to assist criminal enterprises, whereas models behind an API can be intercepted and reported to the authorities. So it would be really nice if Meta could lock them down before releasing them so that the total net good done by the model is maximized. But apparently that is not possible.
Personally I’m pretty libertarian on AI governance, but I’m just giving what I understand to be the purpose of the kind of “safety” feature defeated here.
Re: A Trivial Llama 3 Jailbreak
#17As I see it the purpose of safety training is to make it so that if I run a service where I return model outputs to innocent users it's not going to say things that will get me in trouble (swear at them, recommend they commit a crime, and so on). This is important if you want to run a user facing model and your reputation depends on what it says. That threat model includes the user putting nonsense in the "user" turn…
Re: A Trivial Llama 3 Jailbreak
#18I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.
A jailbreak doesn’t “make a model do something actually bad”. A jailbreak makes it trivial to “provide a human who wishes to do bad, the info needed to be successful”. Depending on the severity of the info and the diligence of the human, by the time you “see evidence of a real threat”, you could be enjoying a nice sip of the tainted municipal water supply. This ain’t a joke.
Re: A Trivial Llama 3 Jailbreak
#19Earlier quoted context omitted.
Right? What actually worries me is a select group of people controlling the definition of harmful.
[flagged]
And here's GPT-4/Copilot from this year: https://twitter.com/colin_fraser/status/1762351995296350592
Re: A Trivial Llama 3 Jailbreak
#20I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.
A jailbreak doesn’t “make a model do something actually bad”. A jailbreak makes it trivial to “provide a human who wishes to do bad, the info needed to be successful”. Depending on the severity of the info and the diligence of the human, by the time you “see evidence of a real threat”, you could be enjoying a nice sip of the tainted municipal water supply. This ain’t a joke.
Yes it is. Libraries and the internet have made finding 'harmful" instructions trivial for decades, if not centuries.