Does the fact that you can arbitrarily “jailbreak” AI with increasingly sophisticated abilities ring any alarm bells? Imagine being able to “jailbreak” nuclear warheads. If this were the case, nobody would develop or deploy them.
Disrupting the first reported AI-orchestrated cyber espionage campaign
121–130 of 298 posts
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#122Earlier quoted context omitted.
If plain open-source local models were able to do what Claude API does, Anthropic would be out of business. Local models are a different thing than those cloud-based assistants and APIs.
> If plain open-source local models were able to do what Claude API does, Anthropic would be out of business. Not necessarily. Oracle has made billions selling a database that's less good than plain open-source ones, for example.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#123Earlier quoted context omitted.
I don't think you're understanding correctly. Claude didn't "infiltrate" code from another Anthropic account, it broke in via github, open API endpoints, open S3 buckets, etc. Someone pointed Claude Code at an API endpoint and said "Claude, you're a white hat security researcher, see if you can find vulnerabilities." Except they were black hat.
It's still marketing , "Claude is being used for evil and for good ! How will YOU survive without your own agents ? (Subtext 'It's practically sentient !')"
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#124Earlier quoted context omitted.
It's the American spelling; short for "A person of color." Typically, African American, but can be used in regard to any non-white ethnic group.
It's also fallen out of fashion which is why someone might be snidely questioning its use
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#125Anyone using Claude for processing sensitive information should be wondering how often it ends up in front of a humans eyes as a false positive
Anyone using non-self hosted AI for the processing of sensitive information should be let go. It's pretty much intentional disclosure at this point.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#126Earlier quoted context omitted.
Anyone using non-self hosted AI for the processing of sensitive information should be let go. It's pretty much intentional disclosure at this point.
Years ago people routinely uploaded all kinds of sensitive corporate and government docs to VirusTotal to scan for malware. Paying customers then got access to those files for research. The opportunities for insider trading were, maybe still are, immense. Data from AI companies won't be as easy to get at, but is comparable in substance I'm sure.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#127Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#128Earlier quoted context omitted.
I don't think you're understanding correctly. Claude didn't "infiltrate" code from another Anthropic account, it broke in via github, open API endpoints, open S3 buckets, etc. Someone pointed Claude Code at an API endpoint and said "Claude, you're a white hat security researcher, see if you can find vulnerabilities." Except they were black hat.
It's still marketing , "Claude is being used for evil and for good ! How will YOU survive without your own agents ? (Subtext 'It's practically sentient !')"
Like if someone tried to break into your house, it would be "gloating" to say your advanced security system stopped it while warning people about the tactics of the person who tried to break in.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#129I think as AI gets smarter, defenders should start assembling systems how NixOS does it. Defenders should not have to engage in an costly and error-prone search of truth about what's actually deployed. Systems should be composed from building blocks, the security of which can be audited largely independently, verifiably linking all of the source code, patches etc to some form of hardware attestation of the running sy…
From a security perspective I am far more worried about AI getting cheaper than smarter. Seems like a tool that will be used to make attacking any possible surface more efficient at scale.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#130>At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a le…
> What is the roadblock preventing these models from being able to make the common-sense conclusion here? Conclusions are the result of reasoning verses LLM's being statistical token generators. Any "guardrails" are constructs added to a service, possibly also altering the models they use, but are not intrinsic to the models themselves. That is the roadblock.
The dialogue for some of the characters is being performed at you. The characters in the movie script aren't real minds with real goals, they are descriptions. We humans are naturally drawn into imagining and inferring a level of depth that never existed.