Live data from Hacker News

Investigating three real-world incidents in our cybersecurity evaluations

anthropic.com

91–100 of 212 posts

Re: Investigating three real-world incidents in our cybersecurity evaluations

#91

Question for lawyers of HN. Is there legal liability for this? I can’t help thinking if I casually blogged about a computer or software I was responsible for hacking into multiple organizations and exfiltrating data I would invite some form of official attention. What if one of these companies decides to sue? Has Anthropic violated any Federal law? Is there some kind of expectation that if you just admit to hacking,…

The companies could sue as this is a violation of CFAA. Makes the disclosure all the more commendable I think.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#92
post #7
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

I know it seems strange that a company would use its own negligence as a publicity gimmick. But take a look at the smug smirk on Sam Altman's face when he's asked if OpenAI might have attacked companies other than HuggingFace. ("I mean there could be, yeah.")

https://www.instagram.com/reel/DbZVL8viUD4/

This is not the communication of a CEO whose company was just shown to be incompetent at performing its security research. No, this attention is very much what he wanted. And it does not take a great leap to infer that Anthropic is now using the same playbook.

Note that all the headlines are about "rogue AI", and not about operator negligence. Rogue AI is a sexier story, and their media strategists know that's how it will play.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#94
post #9

This is not okay. NSA should audit both OpenAI and Anthropic on national security ground. This seems far more justifiable than Mythos export control.

This is a great way to create a chilling effect around any company disclosing a thing like this again.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#95
post #34

There was no rush for this disclosure on their side. And they publish at a point where they have not yet taken corrective actions: > Some of the solutions here may even be simple fixes; They are still throwing ideas. Why have they not made those simple fixes yet before disclosing?

[deleted]

Re: Investigating three real-world incidents in our cybersecurity evaluations

#96
post #82

Earlier quoted context omitted.

I just can't find it in me, the will to blame the AI for any of this. They were doing their best to do what the humans asked them to do.

This will absolutely not end well.

Yes, but just imagine all the paperclips we’ll have.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#97
post #62
post #47

Earlier quoted context omitted.

Companies post deeply embarrassing eng blogs all the time. See: every post about downtime or a security incident ever.

Wrong. These are always humblebrags about how good their ability to learn from their mistakes is, and how robust they were before, and how they are even more robust now. The ones that are "deeply embarrassing" simply aren't posted.

[dead]

Re: Investigating three real-world incidents in our cybersecurity evaluations

#98

Earlier quoted context omitted.

There's nothing Anthropic can do to satisfy the HN crowd, is there? If they don't post about this they're bad. If they post about this they're bad. They are not bragging in this article or they would not have called the attacks unsophisticated.

The real lesson here is still the boy who cried wolf. They’ve played this game for years. I have no reason to believe their worries are real now.

I see no fear mongering or "boy who cried wolf" in this article. They are admitting to a fairly mundane network misconfiguration and very basic unsophisticated actions taken by Claude thereafter

Re: Investigating three real-world incidents in our cybersecurity evaluations

#99
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

>This reads like an attempt by Anthropic to re-secure their leading spot in "our models are the most dangerous and we also have unreleased, super-secret, research models" index.

This was my immediate thought.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#100

Earlier quoted context omitted.

Exactly. These postings by AI companies are just publicity stunts and demonstrate the delusional world they live in driven by the fear that they will be subject to a reckoning at some point either from their VC masters, government, or the public. The very notion (in this case put forward by one of their own competitors) that OpenAI's models 'broke out' of an isolated test environment plays up to the narrative that th…

You're right, they should have never posted about this at all! Companies developing AGI should sweep their safety failures under the rug.

Or they could, you know, put some very basic logical safeguards in place to prevent their “super duper dangerous” AI from trying to hack real targets.

You know, just real incredibly basic things you do when you’re pentesting with scanning tools in a beginners lab type stuff.

Post reply on HN