“distillation attacks” is definitely an interesting way to phrase that.
Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers
111–120 of 382 posts
Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers
#112Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers
#113I haven't been following this story, but the US wanted claude to not be able to find bugs in code?
But then give it exact copy of their house, ask to secure it, which it does and look at what it secured to find out how to get into the original house.
Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers
#114Earlier quoted context omitted.
Ok, and how is that determined? How does anthropic know my "kernel" project isn't a personal toy and not the Linux kernel? How does anthropic determine I'm a legitimate kernel hacker? What proof do I give them and how does it tie back to my email? What would the steps be to create a new project? Do I need to send anthropic a list of my team members each time and keep them updated as the company changes? Shall I be gi…
> What proof do I give them and how does it tie back to my email? Presumably your ID so that feds may pay you a visit when they feel like it, your email need not apply. I’m surprised that there’s even enough pushback against ID verification to matter, all the corpos are probably salivating at the idea of having fully accurate profiles of everyone, think of the ad and product targeting. The govt. would also love that,…
Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers
#115Earlier quoted context omitted.
What's surprising to me is that anyone who has a CS education thinking that jailbreaks are not trivial. It is as simple as normal algorithmic reduction [1], e.g can I transform a dangerous task into a not-dangerous task that the LLM will agree to solve, and then re-transform back. [1]: https://en.wikipedia.org/wiki/Reduction_(complexity)
I think that as simple as is doing a lot of work when the problem domain is all natural language (or more - all strings?) rather than some well specified DSA problem.
For more on this see "Simple Made Easy" by Rich Hickey.
Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers
#116If you set aside political menace, this is a huge problem with Anthropic's strategy. You _cannot_ say that Mythos is super dangerous and can only be rolled out to certain people, but then release Fable with anything other than bulletproof cyber denials. Clearly with LLMs, bulletproof denials are ~impossible due to the way LLMs work. So you've ended up in a situation where Anthropic are simultaneously claiming it's a…
I'm not saying all of Anthropic's statements are true, but mythos did seem to find many legitimate security exploits. You should be able to talk about a helpful-only model being released to limited partners while still releasing a very locked down model that doesn't advance the state of the art on these things, and that seems to be what they did.
There's no inherent contradiction to that.
Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers
#117> “‘Fix this code,’ plus several manual steps to generate test scripts, Feels like the title isn't really giving the full context of what they ended up actually seeing, despite what the lede implies multiple times. Still, ban seems stupid... Still no actual leak of the full "third-party research paper"?
Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers
#118Earlier quoted context omitted.
> What proof do I give them and how does it tie back to my email? Presumably your ID so that feds may pay you a visit when they feel like it, your email need not apply. I’m surprised that there’s even enough pushback against ID verification to matter, all the corpos are probably salivating at the idea of having fully accurate profiles of everyone, think of the ad and product targeting. The govt. would also love that,…
How will the "feds" pay you a visit in Albania or China?
Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers
#119Earlier quoted context omitted.
The idea that an LLM can discern intent on any given prompt is farcical. I might be researching nukes to commit an atrocity, or to prevent one. I might be asking about laundering money to commit a crime, or to prevent one. I might be researching the Nazis because I want to commit a genocide, or I want to read up so I know how to prevent one. Same with cybersecurity. Same with anything. In my opinion, these companies…
I don't really agree with it but the government is moving towards making you ID yourself to use frontier AI - i.e. only US citizens are going to be able to use Claude Fable supposedly. In that regime the AI companies would in fact know if you are a money laundering expert or a normal software engineer. > The idea that an LLM can discern intent on any given prompt is farcical. Not really though. For most people in mos…
And ya, it's pretty easy to hide your intent once you have access.
Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers
#120If you set aside political menace, this is a huge problem with Anthropic's strategy. You _cannot_ say that Mythos is super dangerous and can only be rolled out to certain people, but then release Fable with anything other than bulletproof cyber denials. Clearly with LLMs, bulletproof denials are ~impossible due to the way LLMs work. So you've ended up in a situation where Anthropic are simultaneously claiming it's a…
80 years later, we have something approximating AI, and we're trying to restrict it with simple bright-line rules. Not because we never learned that lesson, but because we simply haven't come up with a better way to do it. Probably because a better way to do it just doesn't exist.
The hilarious part, though, is that it's not the AI that's working around the rules. That's the scenario that's been in science fiction, but it's not what's happening. It's the human users making use of our agency to get the AI agents to work around the rules. Despite calling them "agents", current AI agents don't seem to be able to that particular something. Yet, at least.