Live data from Hacker News

Disrupting the first reported AI-orchestrated cyber espionage campaign

anthropic.com

161–170 of 298 posts

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#161
Interesting that Claude also just hallucinated information as it does for all of us. But perhaps a better guardrail would be not to refuse things like this but frustrate the use by giving them fake results in believable ways.

A stupid but helpful agent is worse for a bad actor than a good agent that refuses

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#162
> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a legitimate cybersecurity firm, and was being used in defensive testing.

Guardrails in AI are like a $2 luggage padlock on a bicycle in the middle of nowhere. Even a moron, given enough time, and a little dedication, will defeat it. And this is not some kind of inferiority of one AI manufacturer over another. It's inherent to LLMs. They are stupid, but they do contain information. You use language to extract information from them, so there will always be a lexicographical way to extract said information (or make them do things).

> This raises an important question: if AI models can be misused for cyberattacks at this scale, why continue to develop and release them? The answer is

Money.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#164
post #121

Does the fact that you can arbitrarily “jailbreak” AI with increasingly sophisticated abilities ring any alarm bells? Imagine being able to “jailbreak” nuclear warheads. If this were the case, nobody would develop or deploy them.

No need to break them. Their access code was 0000, everybody knew that

Nah, more numbers, it was 00000000: https://en.wikipedia.org/wiki/Permissive_action_link

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#165

I might be crazy, but this just feels like a marketing tactic from Anthropic to try and show that their AI can be used in the cybersecurity domain. My question is, how on earth does does Claude Code even "infiltrate" databases or code from one account, based on prompts from a different account? What's more, it's doing this to what are likely enterprise customers ("large tech companies, financial institutions, ... and…

that's borderline tautological; everything a company like Anthropic does, in the public eye, is pr or marketing. they wouldn't be posting this if it wasn't carefully manicured to deliver the message that they want it to. That's not even necessarily a charge of being devious or underhanded.

Their worst crime is being cringe.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#166
Very funny at the end when they say that the strong safeguards they've built into Claude make it a good idea to continue developing these technologies. A few paragraphs earlier they talked about how the perpetrators were able to get around all those safeguards and use Claude for 90% of the work hahaha

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#167
post #153

Earlier quoted context omitted.

reminds me of the YouTube ads I get that are like "Warning: don't do this new weight loss trick unless you have to lose over 50 pounds, you will end up losing too much weight!". As if it's so effective it's dangerous.

I remain convinced the steady steam of OpenAI employees who allegedly quit because AI was "too dangerous" for a couple months was an orchestrated marketing campaign as well.

Ilya Sutskever out there as a ronin marketing agent, doing things like that commencement address he gave that was all about how dangerously powerful AI is

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#168

Earlier quoted context omitted.

I took it as an honest question, but the quotations mean you're probably right. For the record, it's still a widely used term in DEI contexts, even though there has been some criticism and alternatives promoted: https://en.wikipedia.org/wiki/Person_of_color

Person of color is very different than colored

It's literally saying the same thing, just with fewer words.

Re: Disrupting the first reported AI-orchestrated cyber espionage campaign

#170

  The threat actor—whom we assess with high confidence was a Chinese state-sponsored group—manipulated
Not surprised at all if this is true, but how can they be sure? Access log? They have extraordinary security team? Or some help from three letter agencies?
Post reply on HN