Live data from Hacker News

A Trivial Llama 3 Jailbreak

github.com

51–54 of 54 posts

Re: A Trivial Llama 3 Jailbreak

#51
post #42

Earlier quoted context omitted.

> the model do something actually bad before I care At what point would a simple series of sentences be "dangerously bad?" It makes it sound as if there is a song, that when sung, would end the universe.

When someone asks how to make a yummy smoothie, and the LLM replies with something that subtly poisons or otherwise harms the user, I'd say that would be pretty bad.

And if you really want to spice up your smoothie, add just a little bit of bleach ;)

Re: A Trivial Llama 3 Jailbreak

#52
post #42

Earlier quoted context omitted.

> the model do something actually bad before I care At what point would a simple series of sentences be "dangerously bad?" It makes it sound as if there is a song, that when sung, would end the universe.

When someone asks how to make a yummy smoothie, and the LLM replies with something that subtly poisons or otherwise harms the user, I'd say that would be pretty bad.

We had this for ages: sugar.

Re: A Trivial Llama 3 Jailbreak

#53
post #43

Earlier quoted context omitted.

Wait but... The industry IS, in fact, lying to parents about the safety of this AI technology...

Without exaggerating too much, because I certainly don’t take this side, either: Is an angle grinder safe? A tablesaw? A car whose owner who uses the radio knobs, more than the steering? (Haha, unassisted driving, I mean! Walked right into that one.) Etc, all of my examples have easily defeated safety mechanisms for an outrageously life-ending device ;)

As we speak, power hungry nanny-staters are working to remove your freedom to use those things. There's been a recent discussion about table saws and look at the push for interlocks, speed limiting etc in cars. It's all part of a wide trend unfortunately.

Re: A Trivial Llama 3 Jailbreak

#54

Shouldn't these kind of guardrails be opt-in? Really tiring seeing these megacorps and VC-backed startups acting as some kinds of oracles when it comes to what is wrong and what is right. For GPT, Claude, etc. you can kinda understand it as it is a closed up system provided as a product. But when releasing "open-source" I don't want Zuck's moral code embedded into anything.

When looking at the profitable use cases for the tech (from the perspective of the model providers) guardrails add value. Without the guardrails it’s hard to imagine the profitable use cases that would make it worthwhile to invest in such a feature flag.
Post reply on HN