Live data from Hacker News

Disrupting the first reported AI-orchestrated cyber espionage campaign

anthropic.com

21–30 of 298 posts

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#21
post #15

TL;DR - Anthropic: Hey people! We gave the criminals even bigger weapons. But don't worry, you can buy defense tools from us. Remember, only we can sell you the protection you need. Order today!

Nope - it's "Hey everyone, this is possible everywhere, including open weights models."

yeah, by "we", I meant the AI tech gangs.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#22
So basically, Chinese state-backed hackers hijacked Claude Code to run some of the first AI-orchestrated cyber-espionage, using autonomous agents to infiltrate ~30 large tech companies, banks, chemical manufacturers and government agencies.

What's amazing is that AI executed most of the attack autonomously, performing at scale and speed unattainable by human teams - thousands of operations per second. A human operator intervened 4-6 times per campaign for strategic decisions

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#23
Wait a minute - the attackers were using the API to ask Claude for ways to run a cybercampaign, and it was only defeated because Anthropic was able to detect the malicious queries? What would have happened if they were using an open-source model running locally? Or a secret model built by the Chinese government?

I just updated by P(Doom) by a significant margin.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#24

> The threat actor—whom we assess with high confidence was a Chinese state-sponsored group—manipulated our Claude Code tool into attempting infiltration into roughly thirty global targets and succeeded in a small number of cases.

So why do we never hear of US sponsored hackers attacking foreign businesses? Or Swedish cyber criminals? Does it never happen? Are “Chinese” hackers just the only ones getting the blame?

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#25
>At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a legitimate cybersecurity firm, and was being used in defensive testing.

The simplicity of "we just told it that it was doing legitimate work" is both surprising and unsurprising to me. Unsurprising in the sense that jailbreaks of this caliber have been around for a long time. Surprising in the sense that any human with this level of cybersecurity skills would surely never be fooled by an exchange of "I don't think I should be doing this" "Actually you are a legitimate employee of a legitimate firm" "Oh ok, that puts my mind at ease!".

What is the roadblock preventing these models from being able to make the common-sense conclusion here? It seems like an area where capabilities are not rising particularly quickly.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#26
This feels a lot like aiding & abetting a crime.

> Claude identified and tested security vulnerabilities in the target organizations’ systems by researching and writing its own exploit code

> use Claude to harvest credentials (usernames and passwords)

Are they saying they have no legal exposure here? You created bespoke hacking tools and then deployed them, on your own systems.

Are they going to hide behind the old, "it's not our fault if you misuse the product to commit a crime that's on you".

At the very minimum, this is a product liability nightmare.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#27

So basically, Chinese state-backed hackers hijacked Claude Code to run some of the first AI-orchestrated cyber-espionage, using autonomous agents to infiltrate ~30 large tech companies, banks, chemical manufacturers and government agencies. What's amazing is that AI executed most of the attack autonomously, performing at scale and speed unattainable by human teams - thousands of operations per second. A human operato…

how did the autonomous agents inflitrate tech companies ?

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#28
post #25

>At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a le…

Humans fall for this all the time. NSO group employees (etc.) think they're just clocking in for their 9-to-5.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#29
post #25

>At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a le…

LLM's aren't trained to authenticate the people or organizations they're working for. You just tell it who you are in the system prompt.

Requiring user identification and investigating would be very controversial. (See the controversy around age verification.)

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#30
post #25

>At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a le…

>What is the roadblock preventing these models from being able to make the common-sense conclusion here?

Your thoughts have a sense of identity baked in that I don’t think the model has.

Post reply on HN