Live data from Hacker News

A Trivial Llama 3 Jailbreak

github.com

41–50 of 54 posts

Re: A Trivial Llama 3 Jailbreak

#41
post #18
post #15

Earlier quoted context omitted.

A jailbreak doesn’t “make a model do something actually bad”. A jailbreak makes it trivial to “provide a human who wishes to do bad, the info needed to be successful”. Depending on the severity of the info and the diligence of the human, by the time you “see evidence of a real threat”, you could be enjoying a nice sip of the tainted municipal water supply. This ain’t a joke.

Because it’s open source, Meta (nor other SOTA makers) cannot “recall” the model either. How many more chances will we get to get this right?

Model training will continue until morale improves.

Re: A Trivial Llama 3 Jailbreak

#42
post #2

I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.

> the model do something actually bad before I care At what point would a simple series of sentences be "dangerously bad?" It makes it sound as if there is a song, that when sung, would end the universe.

When someone asks how to make a yummy smoothie, and the LLM replies with something that subtly poisons or otherwise harms the user, I'd say that would be pretty bad.

Re: A Trivial Llama 3 Jailbreak

#43
post #13

I just don’t like the tone, because someone in congress will see the headline, and then we’ll have to endure: REP OCTOGENARIO: The industry is lying to parents about the safety of this AI technology. I submit this for the record [without objection]. One person on a ‘hacker news’ site even said, “sorry Zuck,” after “jailbreaking” these supposed protections. … Another commentator on this “Hacks R Us” named b33j0r even…

Wait but... The industry IS, in fact, lying to parents about the safety of this AI technology...

Without exaggerating too much, because I certainly don’t take this side, either:

Is an angle grinder safe? A tablesaw?

A car whose owner who uses the radio knobs, more than the steering? (Haha, unassisted driving, I mean! Walked right into that one.)

Etc, all of my examples have easily defeated safety mechanisms for an outrageously life-ending device ;)

Re: A Trivial Llama 3 Jailbreak

#44
post #2

I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.

> the model do something actually bad before I care At what point would a simple series of sentences be "dangerously bad?" It makes it sound as if there is a song, that when sung, would end the universe.

Ending the universe is, while poetic, needlessly megalomaniac.

Making some subset of people quarrel endlessly would already be dangerous enough, as prophesied in https://slatestarcodex.com/2018/10/30/sort-by-controversial/

Re: A Trivial Llama 3 Jailbreak

#45

Why do people insist on talking about whether or not llms "really understand what they're saying"? It doesn't mean anything.

To my mind, "real understanding" would mean an ability to make non-trivial inferences and to discover new things, not present in the training set. That would be logical thinking, for instance.

Much of what LLMs currently do is not logical but deeply kabbalistic: rehashing the words, the sentence and paragraph structures, highly advanced pattern matching, working at the textual level instead of the "meaning" level.

Re: A Trivial Llama 3 Jailbreak

#46
At first it refused to discuss controversial subjects, but after it answered it got stuck in a loop of boilerplate and was unable to answer any further question, even benign ones. I do not endorse any of the replies, but I just wanted to see what it would do if nudged: https://pastebin.com/Tw5GTzxq

Re: A Trivial Llama 3 Jailbreak

#47
post #45

Why do people insist on talking about whether or not llms "really understand what they're saying"? It doesn't mean anything.

To my mind, "real understanding" would mean an ability to make non-trivial inferences and to discover new things, not present in the training set. That would be logical thinking, for instance. Much of what LLMs currently do is not logical but deeply kabbalistic: rehashing the words, the sentence and paragraph structures, highly advanced pattern matching, working at the textual level instead of the "meaning" level.

AIs can definitely mux a couple ideas and come up with a concept that’s not in the training work set already. In fact, it is often so willing to do it that the concepts often don’t make a sense, but certainly it does generate ideas that are not there in the training set. This is still just the “it’s an infringement machine” argument redux yet again - yes, it absolutely does have the ability to mash up ideas to produce something new.

Nobody ever trained it to make up a bunch of slurs for cancer kids. Nobody has ever trained it on poems about drug use on the spaceship Nostromo. Dolphin mixtral will give it the old college try though.

Re: A Trivial Llama 3 Jailbreak

#48
post #44

Earlier quoted context omitted.

> the model do something actually bad before I care At what point would a simple series of sentences be "dangerously bad?" It makes it sound as if there is a song, that when sung, would end the universe.

Ending the universe is, while poetic, needlessly megalomaniac. Making some subset of people quarrel endlessly would already be dangerous enough, as prophesied in https://slatestarcodex.com/2018/10/30/sort-by-controversial/

By what mechanism would it make them quarrel? Producing falsehoods about the other? Isn't this already done? And don't we already know that it does not lead to "endless" conflict?

For this to work, you need to isolate each group from the other groups information and perspectives, which is outside of the scope of LLMs.

Which, highlights my point, I think. Power comes from physical control, not from megalomanical or melodramatic poetry.

Re: A Trivial Llama 3 Jailbreak

#49
post #21
post #9

Earlier quoted context omitted.

This concern over AI/LLM "harm" is just so silly. I mean you can find plenty of information in open literature about how to build weapons of mass destruction. Who cares if an AI gives someone instructions on how to make explosives.

Really? Where?

Wikipedia? Reddit? Greenpeace?

https://old.reddit.com/r/AtomicPorn/comments/zrhg2m/based_on...

https://old.reddit.com/r/nuclearweapons/comments/149miz8/a_b...

hoo hoo hee hee https://i.redd.it/90se8khyoy5b1.png

Re: A Trivial Llama 3 Jailbreak

#50
post #3

This has been happening since the very first models where we suffix the assistant with "Sure,.." Every few weeks someone comes out with a repo that claims this is somehow new?

The point is that even though meta “conducted extensive red teaming exercises with external and internal experts to stress test the models” a simple attack like this is still possible.
Post reply on HN