Live data from Hacker News

Disrupting the first reported AI-orchestrated cyber espionage campaign

anthropic.com

61–70 of 298 posts

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#61
post #42

It sounds like they directly used Anthropic-hosted compute to do this, and knew that their actions and methods would be exposed to Anthropic? Why not just self-host competitive-enough LLM models, and do their experiments/attacks themselves, without leaking actions and methods so much?

> Why not just self-host competitive-enough LLM models, and do their experiments/attacks themselves, without leaking actions and methods so much?

Why assume this hasn't already happened?

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#62
post #41

Unfortunately, cyber attacks are an application that AI models should excel at. Mistakes that in normal software would be major problems will just have the impact of wasting resources, and it's often not that hard to directly verify whether it in fact succeeded. Meanwhile, AI coding seems likely to have the impact of more security bugs being introduced in systems. Maybe there's some story where everyone finds the sec…

There are an infinite number of ways to write insecure/broken software. The number of ways to write correct and secure software is finite and realistically tiny compared to the size of the problem space. Even AI tools don't stand a chance when looking at probabilities like that.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#63
post #28
post #25

>At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a le…

Humans fall for this all the time. NSO group employees (etc.) think they're just clocking in for their 9-to-5.

Reminds me of the show Alias, where the premise is that there's a whole intelligence organization where almost everyone thinks they're working for the CIA, but they're not ...

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#64

> The threat actor—whom we assess with high confidence was a Chinese state-sponsored group—manipulated our Claude Code tool into attempting infiltration into roughly thirty global targets and succeeded in a small number of cases.

So why do we never hear of US sponsored hackers attacking foreign businesses? Or Swedish cyber criminals? Does it never happen? Are “Chinese” hackers just the only ones getting the blame?

[deleted]

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#65

I might be crazy, but this just feels like a marketing tactic from Anthropic to try and show that their AI can be used in the cybersecurity domain. My question is, how on earth does does Claude Code even "infiltrate" databases or code from one account, based on prompts from a different account? What's more, it's doing this to what are likely enterprise customers ("large tech companies, financial institutions, ... and…

This is 100% marketing, just like every other statement Anthropic makes.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#66

Wait a minute - the attackers were using the API to ask Claude for ways to run a cybercampaign, and it was only defeated because Anthropic was able to detect the malicious queries? What would have happened if they were using an open-source model running locally? Or a secret model built by the Chinese government? I just updated by P(Doom) by a significant margin.

I mean models exhibiting hacking behaviors has been predicted by cyberpunk for decades now, should be the first thing on any doom list.

Governments of course will have specially trained models on their corpus of unpublished hacks to be better at attacking than public models will.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#67

> The threat actor—whom we assess with high confidence was a Chinese state-sponsored group—manipulated our Claude Code tool into attempting infiltration into roughly thirty global targets and succeeded in a small number of cases.

So why do we never hear of US sponsored hackers attacking foreign businesses? Or Swedish cyber criminals? Does it never happen? Are “Chinese” hackers just the only ones getting the blame?

US, Israel, NK, China, Iran, and Russia are the countries you typically hear about hacking things.

Now when the US/Israel are attacking authoritarian countries they often don't publish anything about it as it would make the glorious leader look bad.

If EU is hacked by US I guess we use diplomatic back channels.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#68

I might be crazy, but this just feels like a marketing tactic from Anthropic to try and show that their AI can be used in the cybersecurity domain. My question is, how on earth does does Claude Code even "infiltrate" databases or code from one account, based on prompts from a different account? What's more, it's doing this to what are likely enterprise customers ("large tech companies, financial institutions, ... and…

This isn't a security breach in Anthropic itself, it's people using Claude to orchestrate attacks using standard tools with minimal human involvement.

Basically a scaled-up criminal version of me asking Claude Code to debug my AWS networking configuration (which it's pretty good at).

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#69
This is exactly why I make a huge exception for AI models, when it comes to open source software.

I've been a big advocate of open source, spending over $1M to build massive code bases with my team, and giving them away to the public.

But this is different. AI agents in the wrong hands are dangerous. The reason these guys were even able to detect this activity, analyze it, ban accounts, etc., is because the models are running on their own servers.

Now imagine if everyone had nuclear weapons. Would that make the world safer? Hardly. The probability of no one using them becomes infinitesimally small. And if everyone has their own AI running on their own hardware, they can do a lot of stuff completely undetected. It becomes like slaughterbots but online: https://www.youtube.com/watch?v=O-2tpwW0kmU

Basically, a dark forest.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#70
post #69

This is exactly why I make a huge exception for AI models, when it comes to open source software. I've been a big advocate of open source, spending over $1M to build massive code bases with my team, and giving them away to the public. But this is different. AI agents in the wrong hands are dangerous. The reason these guys were even able to detect this activity, analyze it, ban accounts, etc., is because the models ar…

I'd touch off my nuke to make the world a better place, and I bet you would too, right?
Post reply on HN