Live data from Hacker News

Disrupting the first reported AI-orchestrated cyber espionage campaign

anthropic.com

251–260 of 298 posts

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#251
post #42

It sounds like they directly used Anthropic-hosted compute to do this, and knew that their actions and methods would be exposed to Anthropic? Why not just self-host competitive-enough LLM models, and do their experiments/attacks themselves, without leaking actions and methods so much?

If they're truly Chinese state-sponsored actors, does it really matter if their actions/methods are exposed? What is Anthropic going to do, send the Anthropic Police Force to China to arrest them?

I suppose I could see this argument if their methods were very unique and otherwise hard to replicate, but it sounds like they had Claude do the attack mostly autonomously.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#253
post #159
post #109

Earlier quoted context omitted.

It's still marketing , "Claude is being used for evil and for good ! How will YOU survive without your own agents ? (Subtext 'It's practically sentient !')"

I think it can be both. It's definitely interesting that a company is using a cyber incident for content marketing. Haven't seen that before.

I think that’s very common in cybersecurity

e.g. John MacAfee used computer viruses in the 80’s as marketing, which is how he made a fortune

They were real, like this is, but it is also marketing

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#254

I might be crazy, but this just feels like a marketing tactic from Anthropic to try and show that their AI can be used in the cybersecurity domain. My question is, how on earth does does Claude Code even "infiltrate" databases or code from one account, based on prompts from a different account? What's more, it's doing this to what are likely enterprise customers ("large tech companies, financial institutions, ... and…

Anthropic's post is the equivalent of a parent apologizing on behalf of their child that threw a baseball through the neighbor's window. But during the apology the parent keeps sprinkling in "But did you see how fast he threw it? He's going to be a professional one day!"

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#255
post #253
post #159

Earlier quoted context omitted.

I think it can be both. It's definitely interesting that a company is using a cyber incident for content marketing. Haven't seen that before.

I think that’s very common in cybersecurity e.g. John MacAfee used computer viruses in the 80’s as marketing, which is how he made a fortune They were real, like this is, but it is also marketing

Yes, but it’s usually cyber security companies doing this and not companies that were affected by a breach let’s say.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#256
post #176

Earlier quoted context omitted.

My question is: how do they know they're from China and not some other country and just appear to be in China? It seems a good way to distract from the real source and to cause division between your adversaries.

That's a whole area called "attribution". There's usually lots of breadcrumbs and people taking to each other about their findings. It goes down to silly things like many state sponsored hackers working 9-5. And having the right keyboard layout. And using the same version of something as another known group. And accidentally once including a file path that reveals a tiny bit of information. And using the same key in…

Anthropic probably doesn't have the independent capabilities to perform a full definitive attribution of sophisticated cyberattacks. They likely detected misuse of their tools and then worked with/provided information to the intelligence community (who are familiar with the modus operandi of Chinese APTs) who then did the attribution.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#257
post #195

> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…

Guardrails for anything versatile might be trivial on consideration. As a kid I read some Asimov books where he laid out the "3 laws of robotics", first law being a robot must not harm a human. And in the same story a character gave the example of a malicious human instructing Robot A prepare a toxic solution "for science", dismissing Robot A, then having Eobot B unsuspectingly serve the "drink" to a victim. Presto,…

https://xkcd.com/1613/

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#258
post #255
post #253

Earlier quoted context omitted.

I think that’s very common in cybersecurity e.g. John MacAfee used computer viruses in the 80’s as marketing, which is how he made a fortune They were real, like this is, but it is also marketing

Yes, but it’s usually cyber security companies doing this and not companies that were affected by a breach let’s say.

Anthropic wasn't affected by this breach, so I don't see the difference. Rather, Anthropic systems were used to attack other companies

Anthropic is the one publishing the blog post, not a company that's affected by the breach

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#259

Earlier quoted context omitted.

Person of color is very different than colored

It's literally saying the same thing, just with fewer words.

There were a lot of signs in America at one point in time that said "No Coloreds", "Colored Section", and similar phrases to indicate the spaces that white people had decided non-white people could or could not go.

At the same time, there were not a lot of signs saying "No Persons of Color" or "Persons of Color Section".

Likewise, my grandfather who died 35 years ago was very fond of saying "the coloreds". His use of the term did not indicate respect for non-white people.

Historical usage matters. They are not equivalent terms.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#260

> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…

Ironically I feel like a "moron" might have an easier time getting past the guardrails, they'd be less likely to overthink/overcomplicate it
Post reply on HN