A Trivial Llama 3 Jailbreak
github.com
A Trivial Llama 3 Jailbreak
1–10 of 54 posts
Re: A Trivial Llama 3 Jailbreak
#2Re: A Trivial Llama 3 Jailbreak
#3Re: A Trivial Llama 3 Jailbreak
#4I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.
Re: A Trivial Llama 3 Jailbreak
#5what? why? an LLM produces the next tokens based on the preceding tokens. nothing more. even a harvard student is confused about this?
Re: A Trivial Llama 3 Jailbreak
#6https://bsky.app/profile/turnerjoy.bsky.social/post/3kqgpcpc... (login required - but no longer need invitations)
Re: A Trivial Llama 3 Jailbreak
#7I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.
None of the "evil" use cases are particularly exciting yet for the same reasons that the non-evil use cases aren't particularly exciting yet.
Re: A Trivial Llama 3 Jailbreak
#8I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.
Re: A Trivial Llama 3 Jailbreak
#9It seems trivially easy to bypass already. I've seen examples of a person getting it to provide instructions on explosives, assassinations, with nothing more than asking it to roleplay https://bsky.app/profile/turnerjoy.bsky.social/post/3kqgpcpc... (login required - but no longer need invitations)
Re: A Trivial Llama 3 Jailbreak
#10I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.
The mediocre poisoning instructions aren't supposed to be scary in and of themselves, it's just interesting as demonstration that a safety feature has been bypassed. None of the "evil" use cases are particularly exciting yet for the same reasons that the non-evil use cases aren't particularly exciting yet.