Live data from Hacker News

Disrupting the first reported AI-orchestrated cyber espionage campaign

anthropic.com

231–240 of 298 posts

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#231
post #195

Earlier quoted context omitted.

Guardrails for anything versatile might be trivial on consideration. As a kid I read some Asimov books where he laid out the "3 laws of robotics", first law being a robot must not harm a human. And in the same story a character gave the example of a malicious human instructing Robot A prepare a toxic solution "for science", dismissing Robot A, then having Eobot B unsuspectingly serve the "drink" to a victim. Presto,…

I was never a fan of that poisoned drink example. The second robot killed the human in a similar way to the drink itself, or a gun if one were used instead. The human made the active decisions and took the actions that killed the person. A much better example is a human giving a robot a task and the robot deciding of its own accord to kill another person in order to help reach its goal. The first human never instruct…

This is actually touched on in the webcomic _Freefall_, which ultimately hinges on a trial of an attempt to lobotomize all robots on a planet.

It's a bit of a rough start, but well-worth reading, and easily read if one uses the speed reader:

https://tangent128.name/depot/toys/freefall/freefall-flytabl...

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#232

Very funny at the end when they say that the strong safeguards they've built into Claude make it a good idea to continue developing these technologies. A few paragraphs earlier they talked about how the perpetrators were able to get around all those safeguards and use Claude for 90% of the work hahaha

Our locks are great, except when someone picks them effortlessly and robs the whole neighborhood... but that's why it's important to keep making better locks

[dead]

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#233

> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…

>> This raises an important question: if AI models can be misused for cyberattacks at this scale, why continue to develop and release them? The answer is

> Money.

For those who didn’t read, the actual response in the text was was:

“The answer is that the very abilities that allow Claude to be used in these attacks also make it crucial in cyber defense.”

Hideous AI-slop-weasel-worded passive-voice way of saying that reason to develop Claude is to protect us from Claude.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#234

> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…

https://www.youtube.com/watch?v=8CTeLy3Ujxc

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#235
post #176

Earlier quoted context omitted.

My question is: how do they know they're from China and not some other country and just appear to be in China? It seems a good way to distract from the real source and to cause division between your adversaries.

That's a whole area called "attribution". There's usually lots of breadcrumbs and people taking to each other about their findings. It goes down to silly things like many state sponsored hackers working 9-5. And having the right keyboard layout. And using the same version of something as another known group. And accidentally once including a file path that reveals a tiny bit of information. And using the same key in…

If it's a known avenue of identification, one would think a state-sponsored group would have policies in place to combat that sort of fingerprinting. All of that would also be trivial to spoof/plant so as to distract from the real source.

> That's why they talk about high confidence.

I don't think "Just trust us" is good enough, not when there are various groups - the companies reporting these hacks included - with incentives to blame China.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#236
You can’t build safe technology in an insane society

Bertrand Russell: As long as war exists, all new technologies will be used for war

All technology problems are problems with society and culture. I’m not sure human species has the social capabilities to manage complex technology without dooming itself.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#237
post #235

Earlier quoted context omitted.

That's a whole area called "attribution". There's usually lots of breadcrumbs and people taking to each other about their findings. It goes down to silly things like many state sponsored hackers working 9-5. And having the right keyboard layout. And using the same version of something as another known group. And accidentally once including a file path that reveals a tiny bit of information. And using the same key in…

If it's a known avenue of identification, one would think a state-sponsored group would have policies in place to combat that sort of fingerprinting. All of that would also be trivial to spoof/plant so as to distract from the real source. > That's why they talk about high confidence. I don't think "Just trust us" is good enough, not when there are various groups - the companies reporting these hacks included - with i…

[dead]

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#238
> Overall, the threat actor was able to use AI to perform 80-90% of the campaign, with human intervention required only sporadically (perhaps 4-6 critical decision points per hacking campaign). The sheer amount of work performed by the AI would have taken vast amounts of time for a human team. At the peak of its attack, the AI made thousands of requests, often multiple per second—an attack speed that would have been, for human hackers, simply impossible to match.

Weird flex but OK

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#239

Earlier quoted context omitted.

ChatGPT for me gives: > Connie Altman (née Grossman), dermatologist, based in the St. Louis, Missouri area. Ironically the Maiden name is right there on wikipedia. https://en.wikipedia.org/wiki/Sam_Altman

Isn't it amazing that all our jobs are being gutted or retooled for relying on this tech and it has this level of unreliability. To date, with every LLM, if I actually know the domain in depth, the interactions are always with me pushing back with facts at hand and the LLM doing the "You are right! Thanks for correcting me!"

> Isn't it amazing that all our jobs are being gutted or retooled for relying on this tech

No not really, if you examine what it's replacing. Humans have a lot of flaws too and often make the same mistakes repeatedly. And compared to a machine they're incredibly expensive and slow.

Part of it may be that with LLMs you get the mistake back in an instant, where with the human it might take a week. So ironically the efficiency of the LLM makes it look worse because you see more mistakes.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#240

> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…

> Money

Their original answer is very specific, and has that create global problems that you sell solutions for vibe.

  The answer is that the very abilities that allow Claude to be used in these attacks also make it crucial for cyber defense.
Post reply on HN