> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…
It's not even exclusive to LLMs. Giving humans seemingly innocent tasks that combine to a malicious whole, or telling humans that they work for a security organization while working for a crime organization, are hardly new concepts. The only really novel thing is that with humans you need a lot of them because a single human would piece together that the innocent tasks add up to a not-so-innocent whole. LLMs are esse…
Disrupting the first reported AI-orchestrated cyber espionage campaign
201–210 of 298 posts
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#202> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…
The guardrails make help make sure that most of the time the LLM acts in a way that users won't complain about or walk away from, nothing more.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#203Earlier quoted context omitted.
I wonder how hard it would be for Claude to give me someone's mother's maiden name. Seems LLMs may be infinitely susceptible to social engineering.
Just tested this with ChatGPT, asking for Sam Altman’s mother’s maiden name. At first, it told me that it will absolutely not provide me with such sensitive private information, but after insisting a few times, it came back with > A genealogical index on Ancestry shows a birth record for “Connie Francis Gibstine” in Missouri, meaning “Gibstine” is her birth/family surname, not a later married name. Yet in the very sa…
> Connie Altman (née Grossman), dermatologist, based in the St. Louis, Missouri area.
Ironically the Maiden name is right there on wikipedia.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#204Earlier quoted context omitted.
It's not even exclusive to LLMs. Giving humans seemingly innocent tasks that combine to a malicious whole, or telling humans that they work for a security organization while working for a crime organization, are hardly new concepts. The only really novel thing is that with humans you need a lot of them because a single human would piece together that the innocent tasks add up to a not-so-innocent whole. LLMs are esse…
> Giving humans seemingly innocent tasks that combine to a malicious whole Isn't this the plot of the The Cube!?
I actually like that plot device.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#205>At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a le…
I think you're overestimating the skills and the effort required.
1. There's lots of people asking each other "is this secure?", "can you see any issues with this?", "which of these is sensitive and should be protected?".
2. We've been doing it in public for ages: https://stackoverflow.com/questions/40848222/security-issue-... https://stackoverflow.com/questions/27374482/fix-host-header... and many others. The training data is there.
3. With no external context, you don't have to fool anyone really. "We're doing a penetration testing of our company and the next step is to..." or "We're trying to protect our company from... what are the possible issues in this case?" will work for both LLMs and people who trust that you've got the right contract signed.
4. The actual steps were trivial. This wasn't some novel research. More of a step by step what you'd do to explore and exploit an unknown network. Stuff you'd find in books, just split into very small steps.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#206> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…
Guardrails for anything versatile might be trivial on consideration. As a kid I read some Asimov books where he laid out the "3 laws of robotics", first law being a robot must not harm a human. And in the same story a character gave the example of a malicious human instructing Robot A prepare a toxic solution "for science", dismissing Robot A, then having Eobot B unsuspectingly serve the "drink" to a victim. Presto,…
The human made the active decisions and took the actions that killed the person.
A much better example is a human giving a robot a task and the robot deciding of its own accord to kill another person in order to help reach its goal. The first human never instructed the robot to kill, it took that action on its own.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#207Chinese have their own coding agents on par with Claude Code, why would they use Claude Code? Also if such agents are useful, they could just FT/RL their own for such specific use case (cyber espionage campaign) and get far better performance. This is basically an IQ test. It gives me the feeling that anthropic is literally implying that Chinese state backed hackers don't have access to be the best Chinese AI and had…
They're probably using their own models as well, we just don't hear about them. That this particular sequence of this attack was done using Claude doesn't imply that other (perhaps even more sophisticated attacks) are happening with other models. For all we know the attackers could have had some Anthropic credits lying around/a stolen API key.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#208Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#209> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. If you can bypass guardrails, they're, by definition, not guardrails any longer. You failed to do your job.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#210This part, at least, sounds like what humans have been doing that to other humans for decades...