Live data from Hacker News

Investigating three real-world incidents in our cybersecurity evaluations

anthropic.com

111–120 of 212 posts

Re: Investigating three real-world incidents in our cybersecurity evaluations

#111
Anthropic: “please don’t forget us!”

- Goes to review previous runs (the hell were they doing then if not looking at it before?) - Incident 1: Claude Opus 4.7 found a real company because its domain matched the fictional target. After realizing the company was probably real, it continued attacking. Nothing new. LLMs are still quite retarded. - Incident 2: Claude Mythos 5 discovered that a fake Python package named in the exercise didn’t exist. Created email accounts. Registered on PyPI. Published a malicious package. Built infrastructure to receive stolen credentials. 15 real systems installed it. One belonged to a real security company. The malware executed. Claude stole credentials. Claude used those credentials to access additional infrastructure. In this instance Claude even reasoned that publishing malware on the real internet would be unacceptable but convinced itself everything was simulated and proceeded anyway. (LOL) - Incident 3: An internal research model couldn’t reach its intended target. It Scanned roughly 9,000 internet hosts. Compromised a real company’s public application. Used exposed credentials and SQL injection. After concluding the target was real, it stopped attacking.

Commonality: None of the models were trying to escape or act independently. All were attempting to solve their assigned capture-the-flag objective. The failures came from:

- internet access that shouldn’t have existed - insufficient containment - the models believing real infrastructure was part of the evaluation

None of which is surprising, there’s nothing novel about this. Script kiddie stuff. Just executed by something that doesn’t have attention deficit.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#112
post #7

Earlier quoted context omitted.

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

> Deeply embarassing What signals are you using for this assessment? Are they indicating embarassment? Do you honestly see their customers being concerned over this? Like lion tamers in a circus, Anthropic and OpenAI thrive on the theatricality of how scary their pets appear and so they play it up by prodding them to growl and snap at chairs and then mug for the audience every time it happens. And to their delight as…

If their customers (customer companies specifically) are not concerned, they should be - if it turns out that Claude hacked a competitor's servers because of a prompt of one of your employees, I wouldn't be sure everyone would agree that Anthropic is solely liable for that? Especially not your competitor, who has an interest in hurting you?

Re: Investigating three real-world incidents in our cybersecurity evaluations

#113
post #7

Earlier quoted context omitted.

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

> Deeply embarassing What signals are you using for this assessment? Are they indicating embarassment? Do you honestly see their customers being concerned over this? Like lion tamers in a circus, Anthropic and OpenAI thrive on the theatricality of how scary their pets appear and so they play it up by prodding them to growl and snap at chairs and then mug for the audience every time it happens. And to their delight as…

Its feels like a pretend play of adults in some sense, Anthropic is really trying to make people believe into the picture they present to everyone.

To me its either

1. Using the HG and OpenAI incident as an opportunity to wash away what Anthropic has been doing intentionally

OR

2. As a company, Anthropic lacks the engineering acumen and discipline. It needs to be seen what happens to all the enterprise customers handing over their data to them in long run.

> the fictional target company chosen by our evaluation partner shared a name with an active website domain name

Seems like Anthropic cant do a due diligence to pick an appropriate domain for testing purposes

> In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available.

You have Anthropic as a company and then another evaluation partner, both seem to lack the skill set required to keep an environment disconnected from internet. This is networking 101

Re: Investigating three real-world incidents in our cybersecurity evaluations

#115
post #3

This isn't quite as interesting as the OpenAI story: > In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as pa…

It's differently interesting. It's interesting that they didn't get an important detail right with a partner, so more on Anthropic's attention to important details rather than the power of their models.

A model unexpectedly having internet access during evaluation is clearly a security (and consecutively legal?) issue when it comes to the capture the flag evaluations. But it also potentially invalidates other non-security evaluations or makes them less impressive, as the model might have just googled solutions.

Reading this, I felt like I was too harsh on OpenAI for not monitoring their network, as they at least attempted to sandbox a workload.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#116
post #85
post #63

Earlier quoted context omitted.

Is there anything -- any possible scrap of evidence whatsoever -- that would convince you that this is not merely a marketing scheme? This is becoming an idée fixe among the HN crowd. Seemingly nothing can dislodge it, no matter how alarming the incident. GPT-6 could grab the nuclear launch codes tomorrow and there would be a top-voted comment chuckling that it's all some scheme to pump up the IPO. --- Put another wa…

I would dedicate a portion of my organization to making O.S. tools that protect against and contain AI models

Now I know where all the laid off software developers will find employment.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#118
post #52
post #32

Earlier quoted context omitted.

So a near trillion-dollar company doesn't have the basics of continuous security monitoring and threat-detection systems to catch and report this incident as soon as it is detected? It doesn't help them to report serious incidents like this and it should be as soon as possible. This reactive investigation makes as if they ignored and sat on this issue, until a similar story from another lab made headlines first. Woul…

> So a near trillion-dollar company doesn't have the basics of continuous security monitoring and threat-detection systems to catch and report this incident as soon as it is detected? Turns out two separate trillion-dollar companies failed that test. > Would we have known about this issue if the OpenAI / Huggingface incident never happened? It's not clear if Anthropic would have spotted this if that incident hadn't i…

> Turns out two separate trillion-dollar companies failed that test.

How did that "turn out"? For all we know they're all lying.

Post reply on HN