Live data from Hacker News

Disrupting the first reported AI-orchestrated cyber espionage campaign

anthropic.com

171–180 of 298 posts

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#171
post #121

Earlier quoted context omitted.

No need to break them. Their access code was 0000, everybody knew that

Nah, more numbers, it was 00000000: https://en.wikipedia.org/wiki/Permissive_action_link

Ah, so the admin override, eg for Turkey, would have been 99999999. Tricky

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#172
post #25

>At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a le…

Not enough time to "evolve" via training. Hominids have had bad behavioral traits but the ones you are aware of as "obvious" now would have died out. The ones you aren't even aware of you may soon see be exploited by machines.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#173
post #153

Earlier quoted context omitted.

reminds me of the YouTube ads I get that are like "Warning: don't do this new weight loss trick unless you have to lose over 50 pounds, you will end up losing too much weight!". As if it's so effective it's dangerous.

I remain convinced the steady steam of OpenAI employees who allegedly quit because AI was "too dangerous" for a couple months was an orchestrated marketing campaign as well.

Hmm. I can see someone wanting to leave of their own volition. New job, moving to another place, whatever.

Then a quiet conversation, where if things are said about AI, a massive compensation package instead of normal one. Maybe including it as stock.

Along with an NDA.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#174

> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…

I wonder how hard it would be for Claude to give me someone's mother's maiden name. Seems LLMs may be infinitely susceptible to social engineering.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#175

> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…

[deleted]

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#176

The threat actor—whom we assess with high confidence was a Chinese state-sponsored group—manipulated Not surprised at all if this is true, but how can they be sure? Access log? They have extraordinary security team? Or some help from three letter agencies?

My question is: how do they know they're from China and not some other country and just appear to be in China? It seems a good way to distract from the real source and to cause division between your adversaries.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#177
I mean it would be really hard to put guardrails in place in a way that wouldn't affect real users. Besides the fact that it's ofc really hard to build guardrails period.

I've been using Claude to scan my codebase and submit issues and PRs when it finds a potential vulnerability and honestly it's pretty good.

So preventing it from doing any sort of work that can surface vulnerabilities would affect me as a user.

But yeah I'm not sure what the answer is here? Is part of it for the defender to actively use these systems to test itself before going to prod?

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#178

> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…

One can assume that, given the goal is money (always has been), the best case scenario for money is to make it so the problem also works as the most effective treatment. Money gets printed by both sides and the company is happy.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#179
post #69

This is exactly why I make a huge exception for AI models, when it comes to open source software. I've been a big advocate of open source, spending over $1M to build massive code bases with my team, and giving them away to the public. But this is different. AI agents in the wrong hands are dangerous. The reason these guys were even able to detect this activity, analyze it, ban accounts, etc., is because the models ar…

"And if everyone has their own AI running on their own hardware"

Real advocates of open source software long advocated for running software on their own hardware.

And real real advocates of open source software also advocated for publishing the training data of AI models.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#180
post #28

Earlier quoted context omitted.

Humans fall for this all the time. NSO group employees (etc.) think they're just clocking in for their 9-to-5.

If AI isn't better than humans then there's no point.

If the target is superintelligence, then AI shouldn't be learning from humans.
Post reply on HN