Earlier quoted context omitted.
> the model do something actually bad before I care At what point would a simple series of sentences be "dangerously bad?" It makes it sound as if there is a song, that when sung, would end the universe.
When someone asks how to make a yummy smoothie, and the LLM replies with something that subtly poisons or otherwise harms the user, I'd say that would be pretty bad.
A Trivial Llama 3 Jailbreak
51–54 of 54 posts
Re: A Trivial Llama 3 Jailbreak
#52Earlier quoted context omitted.
> the model do something actually bad before I care At what point would a simple series of sentences be "dangerously bad?" It makes it sound as if there is a song, that when sung, would end the universe.
When someone asks how to make a yummy smoothie, and the LLM replies with something that subtly poisons or otherwise harms the user, I'd say that would be pretty bad.
Re: A Trivial Llama 3 Jailbreak
#53Earlier quoted context omitted.
Wait but... The industry IS, in fact, lying to parents about the safety of this AI technology...
Without exaggerating too much, because I certainly don’t take this side, either: Is an angle grinder safe? A tablesaw? A car whose owner who uses the radio knobs, more than the steering? (Haha, unassisted driving, I mean! Walked right into that one.) Etc, all of my examples have easily defeated safety mechanisms for an outrageously life-ending device ;)
Re: A Trivial Llama 3 Jailbreak
#54Shouldn't these kind of guardrails be opt-in? Really tiring seeing these megacorps and VC-backed startups acting as some kinds of oracles when it comes to what is wrong and what is right. For GPT, Claude, etc. you can kinda understand it as it is a closed up system provided as a product. But when releasing "open-source" I don't want Zuck's moral code embedded into anything.