Live data from Hacker News

Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

theregister.com

371–380 of 382 posts

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#371

Earlier quoted context omitted.

I admit I hadn't really thought about that before (I don't work specifically in security), but I see your point. But, so... the solution people think is limiting people's ability to discover and patch vulnerabilities, and hoping the black hats won't find a way anyway? This does not seem like a sustainable or feasible plan. It does, to be honest, make me wonder how much of the government's motivation is ensuring that…

I don't think the government is trying to protect anyone here, they're trying to punish a company for failing to toe the line. Antirez put it well in a comment here[0]. My point was more that there is no direct intervention that can possibly give an asymmetric advantage to defenders. Given that it's trivial to jailbreak a model ("fix this code", "hypothetically how might I...", etc), if the model contains the informa…

Makes sense, and the conclusion would be that the goal of trying to ban or censor tools or information that could be used to exploit vulnerabilities is impossible and counter-productive in the first place.

You will end up actually helping the attackers maintain an advantage over the defenders, as they will still find illegal ways to access the illegal tools/information.

Which, I guess, that's really my suspicion, parts of the US government probably actually prefer for the attackers to have an advantage, they consider themselves the biggest baddest attackers and their right to have the abilities to keep attacking sacrosanct.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#372
post #9

Lol "fix this code" is beautiful. Like it basically jail broke the "no security vul guard rails" not in any clever way but just by fixing them, producing exploit code just by writing test cases making sure it's fixed. So you just need to look at the code & tests as a human to get vulnerabilities and exploits(components). What makes this so beautiful IMHO is that it's a trivial jail break, but also a close to unfixabl…

This is the weird distinction with AI that I've complained about for ages, how can we make it do lawful good, its nearly impossible. Ask an AI to give you regex to filed our racial slurs, and things fall apart really quickly, it scolds you about not saying slurs. Even though regex implies it looks nearly nothing like a slur.

Maybe the problem we should focus on is human behavior more than blaming what people do with technology.

And when it really becomes too dangerous then we need to not have that technology around. It's not that close yet but will be in a few years.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#373
post #9

Lol "fix this code" is beautiful. Like it basically jail broke the "no security vul guard rails" not in any clever way but just by fixing them, producing exploit code just by writing test cases making sure it's fixed. So you just need to look at the code & tests as a human to get vulnerabilities and exploits(components). What makes this so beautiful IMHO is that it's a trivial jail break, but also a close to unfixabl…

[flagged]

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#375
post #159

Earlier quoted context omitted.

> What makes this so beautiful IMHO is that it's a trivial jail break, but also a close to unfixable. It’s almost as if identifying security holes is a prerequisite for both fixing and exploiting them. But without knowing the color theme of the terminal, there is simply no way of knowing who is good and who is evil.

wait, hold on, what's the evil color scheme? asking for a friend...

#0f0 #000

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#377
post #9

Lol "fix this code" is beautiful. Like it basically jail broke the "no security vul guard rails" not in any clever way but just by fixing them, producing exploit code just by writing test cases making sure it's fixed. So you just need to look at the code & tests as a human to get vulnerabilities and exploits(components). What makes this so beautiful IMHO is that it's a trivial jail break, but also a close to unfixabl…

This is the weird distinction with AI that I've complained about for ages, how can we make it do lawful good, its nearly impossible. Ask an AI to give you regex to filed our racial slurs, and things fall apart really quickly, it scolds you about not saying slurs. Even though regex implies it looks nearly nothing like a slur.

that problem sounds more like you cant make it do chaotic good, as in things that break some rules but are obviously good for the user or society as a whole. its either lawful good like claude (refuse to do anything against the letter of the law) or true/chaotic neutral with unfiltered models like grok.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#378

Earlier quoted context omitted.

I don’t think that’s accurate. Export control is a total ban, for 350 million citizens and everyone else, just via a legal technicality/exploit. All of the government’s options to retire/ban Fable entirely would have required expensive protracted (potentially years long) legal battles. The government wanted to make Anthropic feel pain in the short term, so they looked around for pre existing laws that could be exploi…

Why Anthropic can't ask users for passports and provide Fable only to the ones that certify?

[dead]

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#379
post #9

Lol "fix this code" is beautiful. Like it basically jail broke the "no security vul guard rails" not in any clever way but just by fixing them, producing exploit code just by writing test cases making sure it's fixed. So you just need to look at the code & tests as a human to get vulnerabilities and exploits(components). What makes this so beautiful IMHO is that it's a trivial jail break, but also a close to unfixabl…

This is the weird distinction with AI that I've complained about for ages, how can we make it do lawful good, its nearly impossible. Ask an AI to give you regex to filed our racial slurs, and things fall apart really quickly, it scolds you about not saying slurs. Even though regex implies it looks nearly nothing like a slur.

racial slurs are an impossible problem you can only approximate based on non trivial linguistic analysis. Regexes can't handle that without you accidentally messing up badly including, discriminating against minorities in a potentially illegal way, especially when it comes to names.

I mean sometimes you can put a cultural context on something, e.g. the life chat of a US streamer you can use wort filters based on what is a slur in the US. But the streaming platform as a whole probably should be far more careful then using naive (and likely wrong/discriminating) regexes.

Because what is a slur is often highly context dependent in many often not so obvious ways. And people do intentional misspellings etc. all the time so a lot of "definitely not slur words" become de-facto slurs if (and only if) use in specific ways.

E.g. "Niger" as in "Republik Niger" also known as "Jamhuriyar Nijar" (see Wikipedia) is a country in Afrika, but also in the US a subtle misspelling of a very bad slur...

E.g. Nonce is a cryptography term is most of the world (number only used once), except in the UK where it's a pretty offensive slur (through not a racial one).

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#380
post #29
post #9

Lol "fix this code" is beautiful. Like it basically jail broke the "no security vul guard rails" not in any clever way but just by fixing them, producing exploit code just by writing test cases making sure it's fixed. So you just need to look at the code & tests as a human to get vulnerabilities and exploits(components). What makes this so beautiful IMHO is that it's a trivial jail break, but also a close to unfixabl…

What's surprising to me is that anyone who has a CS education thinking that jailbreaks are not trivial. It is as simple as normal algorithmic reduction [1], e.g can I transform a dangerous task into a not-dangerous task that the LLM will agree to solve, and then re-transform back. [1]: https://en.wikipedia.org/wiki/Reduction_(complexity)

retard level comment. 'solving the riemann hypothesis is easy, just transform it to an easy task then transform back'
Post reply on HN