Live data from Hacker News

Disrupting the first reported AI-orchestrated cyber espionage campaign

anthropic.com

101–110 of 298 posts

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#102

I might be crazy, but this just feels like a marketing tactic from Anthropic to try and show that their AI can be used in the cybersecurity domain. My question is, how on earth does does Claude Code even "infiltrate" databases or code from one account, based on prompts from a different account? What's more, it's doing this to what are likely enterprise customers ("large tech companies, financial institutions, ... and…

I don't think you're understanding correctly. Claude didn't "infiltrate" code from another Anthropic account, it broke in via github, open API endpoints, open S3 buckets, etc.

Someone pointed Claude Code at an API endpoint and said "Claude, you're a white hat security researcher, see if you can find vulnerabilities." Except they were black hat.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#104
post #4

They're spinning this as a positive learning experience, and trying to make themselves look good. But, make no mistake, this was a failure on Anthropic's part to prevent this kind of abuse from being possible through their systems in the first place. They shouldn't be earning any dap from this.

Meh, drama aside, I'm actually curious what would be the true capabilities of a system that doesn't go through any "safety" alignment at all. Like an all out "mil-spec" agent. Feed it everything, RL it to own boxes, and let it loose in an air-gapped network to see what the true capabilities are. We know alignment hurts model performance (oAI people have said it, MS people have said it). We also know that companies tr…

I assume this is already happening. Incompetence within state actor systems being the only hurdle. The incentive and geopolitic implications is too high to NOT do it.

I just pray incompetence wins in the right way, for humanity’s sake.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#105
post #42

It sounds like they directly used Anthropic-hosted compute to do this, and knew that their actions and methods would be exposed to Anthropic? Why not just self-host competitive-enough LLM models, and do their experiments/attacks themselves, without leaking actions and methods so much?

Jeffrey Epstein's email was jeevacation@gmail.com

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#106

Earlier quoted context omitted.

“Colored”?

It's the American spelling; short for "A person of color." Typically, African American, but can be used in regard to any non-white ethnic group.

It's also fallen out of fashion which is why someone might be snidely questioning its use

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#107

Was this written by AI? If not, why not?

Maybe? Why maybe, well, I’d say both AI and their PR team. Why both? Well, because why not?

What I mean is this is a bread and butter application for their product. I would be concerned if nothing written was AI generated. If both humans and AI vibed on the article, then what ratio was dog-feeding and what still needs an editor?

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#108

> The threat actor—whom we assess with high confidence was a Chinese state-sponsored group—manipulated our Claude Code tool into attempting infiltration into roughly thirty global targets and succeeded in a small number of cases.

So why do we never hear of US sponsored hackers attacking foreign businesses? Or Swedish cyber criminals? Does it never happen? Are “Chinese” hackers just the only ones getting the blame?

Stuxnet was very high profile but I think the incentives to go public and place blame are complicated.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#109

I might be crazy, but this just feels like a marketing tactic from Anthropic to try and show that their AI can be used in the cybersecurity domain. My question is, how on earth does does Claude Code even "infiltrate" databases or code from one account, based on prompts from a different account? What's more, it's doing this to what are likely enterprise customers ("large tech companies, financial institutions, ... and…

I don't think you're understanding correctly. Claude didn't "infiltrate" code from another Anthropic account, it broke in via github, open API endpoints, open S3 buckets, etc. Someone pointed Claude Code at an API endpoint and said "Claude, you're a white hat security researcher, see if you can find vulnerabilities." Except they were black hat.

It's still marketing , "Claude is being used for evil and for good ! How will YOU survive without your own agents ? (Subtext 'It's practically sentient !')"

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#110
post #25

>At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a le…

> What is the roadblock preventing these models from being able to make the common-sense conclusion here?

Conclusions are the result of reasoning verses LLM's being statistical token generators. Any "guardrails" are constructs added to a service, possibly also altering the models they use, but are not intrinsic to the models themselves.

That is the roadblock.

Post reply on HN