I might be crazy, but this just feels like a marketing tactic from Anthropic to try and show that their AI can be used in the cybersecurity domain. My question is, how on earth does does Claude Code even "infiltrate" databases or code from one account, based on prompts from a different account? What's more, it's doing this to what are likely enterprise customers ("large tech companies, financial institutions, ... and…
Disrupting the first reported AI-orchestrated cyber espionage campaign
271–280 of 298 posts
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#272> At this point they had to convince Claude—which is extensively trained to avoid harmful behaviors—to engage in the attack. They did so by jailbreaking it, effectively tricking it to bypass its guardrails. They broke down their attacks into small, seemingly innocent tasks that Claude would execute without being provided the full context of their malicious purpose. They also told Claude that it was an employee of a l…
I really think we should stop using the term ‘guard rails’ as it implies a level of control that really doesn’t exist. These things are polite suggestions at best and it’s very misleading to people that do not understand the technology - I’ve got business people saying that using LLMs to process sensitive data is fine because there are “guardrails” in place - we need to make it clear that these kinds of vulnerabiliti…
Think of AI guardrails like the barriers along a highway: they don’t slow the car down, but they do help keep it from veering off course.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#273I think as AI gets smarter, defenders should start assembling systems how NixOS does it. Defenders should not have to engage in an costly and error-prone search of truth about what's actually deployed. Systems should be composed from building blocks, the security of which can be audited largely independently, verifiably linking all of the source code, patches etc to some form of hardware attestation of the running sy…
We soon will have to implement paradoxes in our infrastructure.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#274I think as AI gets smarter, defenders should start assembling systems how NixOS does it. Defenders should not have to engage in an costly and error-prone search of truth about what's actually deployed. Systems should be composed from building blocks, the security of which can be audited largely independently, verifiably linking all of the source code, patches etc to some form of hardware attestation of the running sy…
We soon will have to implement paradoxes in our infrastructure.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#275Earlier quoted context omitted.
It’s not that this is a crazy reach; it’s actually quite a dumb one. Too little pay off, way too much risk. That’s your framework for assessing conspiracies.
Why bring the word “conspiracy” to this discussion though? Marketing stunts aren't conspiracies.
It’s not just a conspiracy, it’s a dumb and harmful one.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#276Earlier quoted context omitted.
> Giving humans seemingly innocent tasks that combine to a malicious whole Isn't this the plot of the The Cube!?
I wouldn't call it the plot of the Cube, more like the setting/world-building.
The construction of the Cube is kind of a backstory, not the main part.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#277Earlier quoted context omitted.
Sorry, your comparative analysis (beyond its rather strange disconnect with your fellow Human beings) ignores the fact that a "stellar" model will fail in this way whereas with us humans, we do get generationally exceptional specimens that push the envelope for the rest of us. To make this crystal clear: Human geniuses were flawed beings but generally you would expect highly reliable utility from their minds. Einstei…
> your comparative analysis (beyond its rather strange disconnect with your fellow Human beings) You seem to be having a different conversataion here. I'm comparing work output by two sources and saying this is why people are choosing to use on over the other for day to day tasks. I'm not waxing poetic about the greater impact to society at large when a new productivity source is introduced. > ignores the fact that a…
Human beings have patterns of behavior that varies from person to person. This is such an established fact that the concept of personal character is a universal and not culturally centered.
(Deterministic) machines and men fail in regular patterns. This is the "human flaws" that you mentioned. It is true that you do not have to be Einstein but the point was missed or not clearly stated. Whether an Einstein or a Joe Random, a person can be observed and we can gauge the capacity of the individual for various tasks. Einstein can be relied upon if we need input on Physics. Random Joe may be an excellent carpenter. Jill writes clearly. Jack is good at organizing people, etc.
So while it is certainly true that human beings are flawed and capabilities are not evenly distributed, they are fairly deterministic components of a production system. Even 'dumb' machines fail in certain characteristic manner, after certain lifetime of service. We know how to make reliable production systems using parts that fail according to patterns.
None of this is true for langauge models and the "AI" built around them. One prompt and your model is "brilliant" and yet entirely possibly it will completely drop the ball in the next sequence. The failure patterns are not deterministic. There is no model, as of now, that would permit the same confidence that we have in building 'fault tolerant systems' using deterministically unreliable/failing parts. None.
Yet every aspect of (cognitive components of) human society is being forcibly affected to incorporate this half-baked technology.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#278Earlier quoted context omitted.
That's a whole area called "attribution". There's usually lots of breadcrumbs and people taking to each other about their findings. It goes down to silly things like many state sponsored hackers working 9-5. And having the right keyboard layout. And using the same version of something as another known group. And accidentally once including a file path that reveals a tiny bit of information. And using the same key in…
If it's a known avenue of identification, one would think a state-sponsored group would have policies in place to combat that sort of fingerprinting. All of that would also be trivial to spoof/plant so as to distract from the real source. > That's why they talk about high confidence. I don't think "Just trust us" is good enough, not when there are various groups - the companies reporting these hacks included - with i…
It relies on people not being perfect and not caring that much. So far, it's working pretty well and the identification leaks are consistent for years.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#279Earlier quoted context omitted.
Anyone using non-self hosted AI for the processing of sensitive information should be let go. It's pretty much intentional disclosure at this point.
Years ago people routinely uploaded all kinds of sensitive corporate and government docs to VirusTotal to scan for malware. Paying customers then got access to those files for research. The opportunities for insider trading were, maybe still are, immense. Data from AI companies won't be as easy to get at, but is comparable in substance I'm sure.
Re: Disrupting the first reported AI-orchestrated cyber espionage campaign
#280Earlier quoted context omitted.
If it's a known avenue of identification, one would think a state-sponsored group would have policies in place to combat that sort of fingerprinting. All of that would also be trivial to spoof/plant so as to distract from the real source. > That's why they talk about high confidence. I don't think "Just trust us" is good enough, not when there are various groups - the companies reporting these hacks included - with i…
> If it's a known avenue of identification, one would think a state-sponsored group would have policies in place to combat that sort of fingerprinting. It relies on people not being perfect and not caring that much. So far, it's working pretty well and the identification leaks are consistent for years.