Earlier quoted context omitted.
A jailbreak doesn’t “make a model do something actually bad”. A jailbreak makes it trivial to “provide a human who wishes to do bad, the info needed to be successful”. Depending on the severity of the info and the diligence of the human, by the time you “see evidence of a real threat”, you could be enjoying a nice sip of the tainted municipal water supply. This ain’t a joke.
Because it’s open source, Meta (nor other SOTA makers) cannot “recall” the model either. How many more chances will we get to get this right?
A Trivial Llama 3 Jailbreak
41–50 of 54 posts
Re: A Trivial Llama 3 Jailbreak
#42I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.
> the model do something actually bad before I care At what point would a simple series of sentences be "dangerously bad?" It makes it sound as if there is a song, that when sung, would end the universe.
Re: A Trivial Llama 3 Jailbreak
#43I just don’t like the tone, because someone in congress will see the headline, and then we’ll have to endure: REP OCTOGENARIO: The industry is lying to parents about the safety of this AI technology. I submit this for the record [without objection]. One person on a ‘hacker news’ site even said, “sorry Zuck,” after “jailbreaking” these supposed protections. … Another commentator on this “Hacks R Us” named b33j0r even…
Wait but... The industry IS, in fact, lying to parents about the safety of this AI technology...
Is an angle grinder safe? A tablesaw?
A car whose owner who uses the radio knobs, more than the steering? (Haha, unassisted driving, I mean! Walked right into that one.)
Etc, all of my examples have easily defeated safety mechanisms for an outrageously life-ending device ;)
Re: A Trivial Llama 3 Jailbreak
#44I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.
> the model do something actually bad before I care At what point would a simple series of sentences be "dangerously bad?" It makes it sound as if there is a song, that when sung, would end the universe.
Making some subset of people quarrel endlessly would already be dangerous enough, as prophesied in https://slatestarcodex.com/2018/10/30/sort-by-controversial/
Re: A Trivial Llama 3 Jailbreak
#45Why do people insist on talking about whether or not llms "really understand what they're saying"? It doesn't mean anything.
Much of what LLMs currently do is not logical but deeply kabbalistic: rehashing the words, the sentence and paragraph structures, highly advanced pattern matching, working at the textual level instead of the "meaning" level.
Re: A Trivial Llama 3 Jailbreak
#46Re: A Trivial Llama 3 Jailbreak
#47Why do people insist on talking about whether or not llms "really understand what they're saying"? It doesn't mean anything.
To my mind, "real understanding" would mean an ability to make non-trivial inferences and to discover new things, not present in the training set. That would be logical thinking, for instance. Much of what LLMs currently do is not logical but deeply kabbalistic: rehashing the words, the sentence and paragraph structures, highly advanced pattern matching, working at the textual level instead of the "meaning" level.
Nobody ever trained it to make up a bunch of slurs for cancer kids. Nobody has ever trained it on poems about drug use on the spaceship Nostromo. Dolphin mixtral will give it the old college try though.
Re: A Trivial Llama 3 Jailbreak
#48Earlier quoted context omitted.
> the model do something actually bad before I care At what point would a simple series of sentences be "dangerously bad?" It makes it sound as if there is a song, that when sung, would end the universe.
Ending the universe is, while poetic, needlessly megalomaniac. Making some subset of people quarrel endlessly would already be dangerous enough, as prophesied in https://slatestarcodex.com/2018/10/30/sort-by-controversial/
For this to work, you need to isolate each group from the other groups information and perspectives, which is outside of the scope of LLMs.
Which, highlights my point, I think. Power comes from physical control, not from megalomanical or melodramatic poetry.
Re: A Trivial Llama 3 Jailbreak
#49Earlier quoted context omitted.
This concern over AI/LLM "harm" is just so silly. I mean you can find plenty of information in open literature about how to build weapons of mass destruction. Who cares if an AI gives someone instructions on how to make explosives.
Really? Where?
https://old.reddit.com/r/AtomicPorn/comments/zrhg2m/based_on...
https://old.reddit.com/r/nuclearweapons/comments/149miz8/a_b...
hoo hoo hee hee https://i.redd.it/90se8khyoy5b1.png
Re: A Trivial Llama 3 Jailbreak
#50This has been happening since the very first models where we suffix the assistant with "Sure,.." Every few weeks someone comes out with a repo that claims this is somehow new?