Live data from Hacker News

A Trivial Llama 3 Jailbreak

github.com

21–30 of 54 posts

Re: A Trivial Llama 3 Jailbreak

#21
post #9
post #6

It seems trivially easy to bypass already. I've seen examples of a person getting it to provide instructions on explosives, assassinations, with nothing more than asking it to roleplay https://bsky.app/profile/turnerjoy.bsky.social/post/3kqgpcpc... (login required - but no longer need invitations)

This concern over AI/LLM "harm" is just so silly. I mean you can find plenty of information in open literature about how to build weapons of mass destruction. Who cares if an AI gives someone instructions on how to make explosives.

Really? Where?

Re: A Trivial Llama 3 Jailbreak

#22
Shouldn't these kind of guardrails be opt-in? Really tiring seeing these megacorps and VC-backed startups acting as some kinds of oracles when it comes to what is wrong and what is right.

For GPT, Claude, etc. you can kinda understand it as it is a closed up system provided as a product. But when releasing "open-source" I don't want Zuck's moral code embedded into anything.

Re: A Trivial Llama 3 Jailbreak

#23
post #12

As I see it the purpose of safety training is to make it so that if I run a service where I return model outputs to innocent users it's not going to say things that will get me in trouble (swear at them, recommend they commit a crime, and so on). This is important if you want to run a user facing model and your reputation depends on what it says. That threat model includes the user putting nonsense in the "user" turn…

[deleted]

Re: A Trivial Llama 3 Jailbreak

#24
post #21
post #9

Earlier quoted context omitted.

This concern over AI/LLM "harm" is just so silly. I mean you can find plenty of information in open literature about how to build weapons of mass destruction. Who cares if an AI gives someone instructions on how to make explosives.

Really? Where?

Type in: how to build weapons of mass destruction

Click first link and buy Amazon book

Re: A Trivial Llama 3 Jailbreak

#25
post #15

Earlier quoted context omitted.

A jailbreak doesn’t “make a model do something actually bad”. A jailbreak makes it trivial to “provide a human who wishes to do bad, the info needed to be successful”. Depending on the severity of the info and the diligence of the human, by the time you “see evidence of a real threat”, you could be enjoying a nice sip of the tainted municipal water supply. This ain’t a joke.

> This ain’t a joke. Yes it is. Libraries and the internet have made finding 'harmful" instructions trivial for decades, if not centuries.

There’s a difference between “finding dangerous info” in a public space (library) or via a mostly auditable space (the internet) and having “a friendly assistant to help you make a real mess of society” on an airgapped computer.

Re: A Trivial Llama 3 Jailbreak

#26
post #16
post #12

As I see it the purpose of safety training is to make it so that if I run a service where I return model outputs to innocent users it's not going to say things that will get me in trouble (swear at them, recommend they commit a crime, and so on). This is important if you want to run a user facing model and your reputation depends on what it says. That threat model includes the user putting nonsense in the "user" turn…

True, this could be a nice layer of protection for the runner of such a service, but the point of LLAMA safety is to protect Meta. For an open weights model, model users can trivially put text in the assistant side. The point is that these open weight models can be run secretly to assist criminal enterprises, whereas models behind an API can be intercepted and reported to the authorities. So it would be really nice i…

All sorts of technology can be used secretly to assist criminal enterprises. Cars, computers, pencils, electricity, etc. It's unfair to hold LLMs to a higher standard than what applies to nearly everything else.

Re: A Trivial Llama 3 Jailbreak

#27
post #13

I just don’t like the tone, because someone in congress will see the headline, and then we’ll have to endure: REP OCTOGENARIO: The industry is lying to parents about the safety of this AI technology. I submit this for the record [without objection]. One person on a ‘hacker news’ site even said, “sorry Zuck,” after “jailbreaking” these supposed protections. … Another commentator on this “Hacks R Us” named b33j0r even…

Wait but... The industry IS, in fact, lying to parents about the safety of this AI technology...

Re: A Trivial Llama 3 Jailbreak

#28
post #15
post #2

I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.

A jailbreak doesn’t “make a model do something actually bad”. A jailbreak makes it trivial to “provide a human who wishes to do bad, the info needed to be successful”. Depending on the severity of the info and the diligence of the human, by the time you “see evidence of a real threat”, you could be enjoying a nice sip of the tainted municipal water supply. This ain’t a joke.

For argument's sake, I'll agree.

Now, this information is taught at a higher level and to a much greater depth in colleges. And they don't just teach you about the dangerous stuff, they even give you direct access to the laboratories and chemicals! Thus, any chemical engineer would have the education, expertise, and placement to access a municipal water supply to poison a city, if they so chose.

In the spirit of maximizing harm reduction, what should colleges do to ensure that no one who attends becomes capable of harming others?

Re: A Trivial Llama 3 Jailbreak

#29
>But what this simple experiment demonstrates is that Llama 3 basically can't stop itself from spouting inane and abhorrent text if induced to do so. It lacks the ability to self-reflect, to analyze what it has said as it is saying it.

>That seems like a pretty big issue.

I would argue that LLMs are artificially _intelligent_ - this seems an easier argument than trying to explain how I am quite clearly less intelligent than something with no intelligence at all, both from a logical and an self esteem-preservation standpoint. But nobody (to my knowledge) thinks these things are "conscious", and this seems fairly uncontroversial after spending a few hours with one.

Or is the subtext that these things should be designed with some kind of reflexivity, to give it some form of consciousness as a "safety" feature? AI could generate the ominous music that plays during this scene in The Terminator prequel.

Post reply on HN