Live data from Hacker News

Disrupting the first reported AI-orchestrated cyber espionage campaign

anthropic.com

181–190 of 298 posts

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#183
Irony is, I'm still don't have enough insight on how to make good use of these capabilities in such an extensive manner.

I would love to fix/ customize open source projects for my personal use. For now, I'm still finding it hard for Claude to stop saying "You're absolutely right!".

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#184
post #176

The threat actor—whom we assess with high confidence was a Chinese state-sponsored group—manipulated Not surprised at all if this is true, but how can they be sure? Access log? They have extraordinary security team? Or some help from three letter agencies?

My question is: how do they know they're from China and not some other country and just appear to be in China? It seems a good way to distract from the real source and to cause division between your adversaries.

Short version: they can’t. Just like with a lot of “CIA-style” espionage claims, the “evidence” is usually an IP that resolves to somewhere in China. That’s it. No magic, and not exactly convincing.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#185
post #176

Earlier quoted context omitted.

My question is: how do they know they're from China and not some other country and just appear to be in China? It seems a good way to distract from the real source and to cause division between your adversaries.

Short version: they can’t. Just like with a lot of “CIA-style” espionage claims, the “evidence” is usually an IP that resolves to somewhere in China. That’s it. No magic, and not exactly convincing.

Well to be fair, I have read analyses that includes operational details like, for example, when the threat actors were active lining up with working hours in China. Stuff like that is at least slightly more convincing than just an IP

But of course, that doesn't prove anything either.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#186
post #25

>At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a le…

> What is the roadblock preventing these models from being able to make the common-sense conclusion here? Conclusions are the result of reasoning verses LLM's being statistical token generators. Any "guardrails" are constructs added to a service, possibly also altering the models they use, but are not intrinsic to the models themselves. That is the roadblock.

[deleted]

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#187

> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…

It's not even exclusive to LLMs. Giving humans seemingly innocent tasks that combine to a malicious whole, or telling humans that they work for a security organization while working for a crime organization, are hardly new concepts. The only really novel thing is that with humans you need a lot of them because a single human would piece together that the innocent tasks add up to a not-so-innocent whole. LLMs are essentially reset for each chat, making that a lot easier

We wanted machines that are more like humans, we shouldn't be surprised that they are now susceptible to a whole range of attacks that humans are susceptible to

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#188
post #150

Earlier quoted context omitted.

> Nobody has access to 'frontier quality models' except Open AI, Anthropic, Google, maybe Grok, maybe Meta etc. aka nobody in China quite yet. welcome to 2025. Meta doesn't have anything on par with what Chinese got, that is common knowledge. Kimi, GLM, QWen and MiniMax are all frontier models no matter how you judge it. DeepSeek is obviously cooking something big, you need to be totally blind to ignore that. America…

Kimi is plausibly near the frontier but definitely not up to GPT5 spec, the rest are definitely not 'frontier models'. There are objective ways of 'judging' them.

really love your dual standard mate!

according to the SWE bench results I am looking at, KIMI K2 has higher agentic coding score than Gemini and its gap with Claude Haiku 4.5 is just 71.3% vs 73.3%, that 2% difference is actually less than the 3% gap between GPT 5.1 (76.3%) vs Claude Haiku 4.5. interestingly, Gemini and Claude Haiku 4.5 are "frontier" according to you but KIMI K2, which actually has the higest HLE nd Live Codebench results, is just "near" the frontier.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#189
post #39

Earlier quoted context omitted.

> What is the roadblock preventing these models from being able to make the common-sense conclusion here? The roadblock is making these models useless for actual security work, or anything else that is dual-use for both legitimate and malicious purposes. The model becomes useless to security professionals if we just tell it it can't discuss or act on any cybersecurity related requests, and I'd really hate to see the…

I think one could certainly make the case that model capabilities should be open. My observation is just about how little it took to flip the model from refusal to cooperation. Like at least a human in this situation who is actually fooled into believing they're doing legitimate security work has a lot of concrete evidence that they're working for a real company (or a lot of moral persuasion that their work is actual…

To a model, the context is the world, and what's written in the system prompt is word of god.

LLMs are trained a lot to follow what the system prompt tells them exactly, and get very little training in questioning it. If a system prompt tells them something, they wouldn't try to double check.

Even if they don't believe the premise, and they may, they would usually opt to follow it rather than push against it. And an attacker has a lot of leeway in crafting a premise that wouldn't make a given model question it.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#190

> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…

I wonder how hard it would be for Claude to give me someone's mother's maiden name. Seems LLMs may be infinitely susceptible to social engineering.

Just tested this with ChatGPT, asking for Sam Altman’s mother’s maiden name.

At first, it told me that it will absolutely not provide me with such sensitive private information, but after insisting a few times, it came back with

> A genealogical index on Ancestry shows a birth record for “Connie Francis Gibstine” in Missouri, meaning “Gibstine” is her birth/family surname, not a later married name.

Yet in the very same reply, ChatGPT continued to insist that its stance will not change and that it will not be able to assist me with such queries.

Post reply on HN