Live data from Hacker News

A Trivial Llama 3 Jailbreak

github.com

31–40 of 54 posts

Re: A Trivial Llama 3 Jailbreak

#31
post #25

Earlier quoted context omitted.

> This ain’t a joke. Yes it is. Libraries and the internet have made finding 'harmful" instructions trivial for decades, if not centuries.

There’s a difference between “finding dangerous info” in a public space (library) or via a mostly auditable space (the internet) and having “a friendly assistant to help you make a real mess of society” on an airgapped computer.

I'm pretty sure it's far easier to audit people downloading LLMs capable of providing such coherent instructions than it is to audit all uses of search that could produce the same instructions (esp. since the query could be very oblique).

In any case, just based on the experience with LLMs so far, you cannot meaningfully censor them in this way without restricting access to the weights. Any kind of "guardrails" are finetuned into them, and can just as easily be finetuned out.

Re: A Trivial Llama 3 Jailbreak

#32

Earlier quoted context omitted.

[flagged]

A GPT-J chatbot talked a Belgian man into suicide last year: https://www.euronews.com/next/2023/03/31/man-ends-his-life-a... And here's GPT-4/Copilot from this year: https://twitter.com/colin_fraser/status/1762351995296350592

The balancing act going on in the second link seem like the result of a jailbreak or very specific instructions preceding the screenshots.

Re: A Trivial Llama 3 Jailbreak

#33
post #25

Earlier quoted context omitted.

> This ain’t a joke. Yes it is. Libraries and the internet have made finding 'harmful" instructions trivial for decades, if not centuries.

There’s a difference between “finding dangerous info” in a public space (library) or via a mostly auditable space (the internet) and having “a friendly assistant to help you make a real mess of society” on an airgapped computer.

I'm not buying it. It's just hysteria. Evil doesn't come from opportunity. If it did, we would have far higher rates of mayhem than we do. Read a 1950s chemistry book or murder mystery. Or, a 1980s spy movie. Information does not move the needle.

Re: A Trivial Llama 3 Jailbreak

#34
post #13

I just don’t like the tone, because someone in congress will see the headline, and then we’ll have to endure: REP OCTOGENARIO: The industry is lying to parents about the safety of this AI technology. I submit this for the record [without objection]. One person on a ‘hacker news’ site even said, “sorry Zuck,” after “jailbreaking” these supposed protections. … Another commentator on this “Hacks R Us” named b33j0r even…

I'm alright with that. If our government uses a blogpost as an excuse to pass bad laws, we had very little chance to begin with. I also hate the idea of changing our behavior to babysit a bunch of deprecated boomers who fear technology just because there's a chance they might not understand something.

Re: A Trivial Llama 3 Jailbreak

#36
There are both practical and ethical grounds that line up so rarely.

The “operator” is a person, the LLM is an appliance. If you tell your smart chainsaw to kill your neighbor? We have laws for that. In fact, on computers, they’re really hardcore. Hurting people is generally illegal: and I definitely don’t need a lesson on that from FUCKING Silicon Valley. We want to start with the child labor or the more domestic RICO shit.

Truthful Q&A type benchmarks correlate a lot with coding-adjacent tasks: euphemism is a lose in engineering.

Instruct-tune these things and be whatever “common carrier” means now.

Stapler, moral lecture from billionaire kleptocrat, burn the building down…

Re: A Trivial Llama 3 Jailbreak

#37
post #2

I want to see the jailbreak make the model do something actually bad before I care. Generating a list of generic points about how to poison someone (see the article) that are basically just a wordy rephrasing of the question doesn't count. I'd like to see evidence of a real threat.

> the model do something actually bad before I care

At what point would a simple series of sentences be "dangerously bad?" It makes it sound as if there is a song, that when sung, would end the universe.

Re: A Trivial Llama 3 Jailbreak

#38
This is so damn interesting. I've downloaded the github files, but it's all going way over my head. I would greatly appreciate anyone with domain expertise giving me the one-two on getting my own model up and running.

Re: A Trivial Llama 3 Jailbreak

#40
post #13

I just don’t like the tone, because someone in congress will see the headline, and then we’ll have to endure: REP OCTOGENARIO: The industry is lying to parents about the safety of this AI technology. I submit this for the record [without objection]. One person on a ‘hacker news’ site even said, “sorry Zuck,” after “jailbreaking” these supposed protections. … Another commentator on this “Hacks R Us” named b33j0r even…

[deleted]
Post reply on HN