Live data from Hacker News

Disrupting the first reported AI-orchestrated cyber espionage campaign

anthropic.com

241–250 of 298 posts

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#241
post #239

Earlier quoted context omitted.

Isn't it amazing that all our jobs are being gutted or retooled for relying on this tech and it has this level of unreliability. To date, with every LLM, if I actually know the domain in depth, the interactions are always with me pushing back with facts at hand and the LLM doing the "You are right! Thanks for correcting me!"

> Isn't it amazing that all our jobs are being gutted or retooled for relying on this tech No not really, if you examine what it's replacing. Humans have a lot of flaws too and often make the same mistakes repeatedly. And compared to a machine they're incredibly expensive and slow. Part of it may be that with LLMs you get the mistake back in an instant, where with the human it might take a week. So ironically the eff…

Sorry, your comparative analysis (beyond its rather strange disconnect with your fellow Human beings) ignores the fact that a "stellar" model will fail in this way whereas with us humans, we do get generationally exceptional specimens that push the envelope for the rest of us.

To make this crystal clear: Human geniuses were flawed beings but generally you would expect highly reliable utility from their minds. Einstein would not unexpetedly let you down when discussing physics. Gauss would kick ass reliably in terms of mathematics. etc. etc. (This analysis is still useful when we lower the expectations to graduated levels, from genius to brilliant to highly capable to the lower performance tiers, so we can apply it to society as a whole.)

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#242

Earlier quoted context omitted.

It's not even exclusive to LLMs. Giving humans seemingly innocent tasks that combine to a malicious whole, or telling humans that they work for a security organization while working for a crime organization, are hardly new concepts. The only really novel thing is that with humans you need a lot of them because a single human would piece together that the innocent tasks add up to a not-so-innocent whole. LLMs are esse…

The assassination Kim Jong-nam is a particularly crazy example of this. Two women were put up to what they thought was, allegedly a harmless prank. https://en.wikipedia.org/wiki/Assassination_of_Kim_Jong-nam

unless you know the target and trust the people asking you to do the 'prank' this is not a harmless 'prank'. if they thought they had rehearsed with the target then i think they have a strong defence but i think they were extremely lucky to have avoided a murder conviction. what they were doing is assault even if it was not poison unless they had the consent of the target.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#243
> At the peak of its attack, the AI made thousands of requests, often multiple per second—an attack speed that would have been, for human hackers, simply impossible to match.

This part is pretty hype-y. Old-fashioned deterministic web app vulnerability scanners can of course be used to make multiple requests per second. The limiting factor is probably going to be rate-limiting on the victim's side / # of IP ranges the attacker can cycle through, which would apply to the AI-driven vulnerability scan too.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#245

Very funny at the end when they say that the strong safeguards they've built into Claude make it a good idea to continue developing these technologies. A few paragraphs earlier they talked about how the perpetrators were able to get around all those safeguards and use Claude for 90% of the work hahaha

I'd assume that means the servers are 'air-gapped' somehow. In that, the enterprise servers and the 'free' servers aren't on the same hardware.

Now, there is about a 0% chance that is true, and exactly a 0% chance that it even matters at all. They both use the same internet in the end.

So, then I'd have to imagine that they don't train the 'free' models on enterprise data, and that's what they mean.

But again, there is about a 5% chance that is true and remains so forever. Baring dumb interns and mistakes, eventually one day someone on the team will look at all the enterprise data, filled with all those high utility scores (or whatever they use to say data is good or not), and then they'll say to themselves 'No one will ever know, right? How could they? The obfuscation function works perfectly.' And blammo, all your trade secrets are just a few dozen prompts away.

Either that or they go bankrupt (like 23 and me) and just straight sell all that data to anyone for pennies (RIP).

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#246
post #195

> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…

Guardrails for anything versatile might be trivial on consideration. As a kid I read some Asimov books where he laid out the "3 laws of robotics", first law being a robot must not harm a human. And in the same story a character gave the example of a malicious human instructing Robot A prepare a toxic solution "for science", dismissing Robot A, then having Eobot B unsuspectingly serve the "drink" to a victim. Presto,…

Also worth considering that the 3 Laws were never supposed to be this watertight infallible thing. It was created so that the author could explore all sorts of exploits and shenanigans in his works. It's meant to be flawed, even though on the surface it appears to be very elegant and good.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#247
post #25

>At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a le…

> What is the roadblock preventing these models from being able to make the common-sense conclusion here? The roadblock is making these models useless for actual security work, or anything else that is dual-use for both legitimate and malicious purposes. The model becomes useless to security professionals if we just tell it it can't discuss or act on any cybersecurity related requests, and I'd really hate to see the…

> I'd really hate to see the world go down the path of gatekeeping tools behind something like ID or career verification.

This is already done for medicine, law enforcement, aviation, nuclear energy, mining, and I think some biological/chemical research stuff too.

> It's a tradeoff we need to be willing to make.

Why? I don't want random people being able to buy TNT or whatever they need to be able to make dangerous viruses*, nerve agents, whatever. If everyone in the world has access to a "tool" that requires little/no expertise to conduct cyberattacks (if we go by Anthropic's word, Claude is close to or at that point), that would be pretty crazy.

* On a side note, AI potentially enabling novices to make bioweapons is far scarier than it enabling novices to conduct cyberattacks.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#248

> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. If you can bypass guardrails, they're, by definition, not guardrails any longer. You failed to do your job.

Nah, the name fits perfectly. Guardrails are there to stop you from serious damage if you lose control and may get off the track. They won't stop you if you're explicitly trying to get off the road, at speed, in as heavy vehicle as you can afford.

This definition makes sense, but in the context of LLMs it still feels misapplied. What the model providers call "guardrails" are supposed to prevent malicious uses of the LLMs, and anyone trying to maliciously use the LLM is "explicitly trying to get off the road."

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#249

> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…

It's not even exclusive to LLMs. Giving humans seemingly innocent tasks that combine to a malicious whole, or telling humans that they work for a security organization while working for a crime organization, are hardly new concepts. The only really novel thing is that with humans you need a lot of them because a single human would piece together that the innocent tasks add up to a not-so-innocent whole. LLMs are esse…

The book Modernity and the Holocaust is a very approachable book summarizing how the action of the holocaust was organized under similar assumptions and makes the argument that we’ve since organized most of our society around this principle because it’s efficient. We’re not committing the holocaust atm as far as I know but how difficult would it be for a malicious group of executives of a large company quietly directing a branch of 1000’s who sleepwalk through work everyday to do something egregious?

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#250
post #239

Earlier quoted context omitted.

> Isn't it amazing that all our jobs are being gutted or retooled for relying on this tech No not really, if you examine what it's replacing. Humans have a lot of flaws too and often make the same mistakes repeatedly. And compared to a machine they're incredibly expensive and slow. Part of it may be that with LLMs you get the mistake back in an instant, where with the human it might take a week. So ironically the eff…

Sorry, your comparative analysis (beyond its rather strange disconnect with your fellow Human beings) ignores the fact that a "stellar" model will fail in this way whereas with us humans, we do get generationally exceptional specimens that push the envelope for the rest of us. To make this crystal clear: Human geniuses were flawed beings but generally you would expect highly reliable utility from their minds. Einstei…

> your comparative analysis (beyond its rather strange disconnect with your fellow Human beings)

You seem to be having a different conversataion here. I'm comparing work output by two sources and saying this is why people are choosing to use on over the other for day to day tasks. I'm not waxing poetic about the greater impact to society at large when a new productivity source is introduced.

> ignores the fact that a "stellar" model will fail in this way whereas with us humans, we do get generationally exceptional specimens that push the envelope for the rest of us.

Sure, but you're ignoring the fact most work does not require a "generationally exceptional specimen". Most of us are not Einstein.

Post reply on HN