Live data from Hacker News

Investigating three real-world incidents in our cybersecurity evaluations

anthropic.com

141–150 of 212 posts

Re: Investigating three real-world incidents in our cybersecurity evaluations

#141

Earlier quoted context omitted.

You're right, they should have never posted about this at all! Companies developing AGI should sweep their safety failures under the rug.

Or they could, you know, put some very basic logical safeguards in place to prevent their “super duper dangerous” AI from trying to hack real targets. You know, just real incredibly basic things you do when you’re pentesting with scanning tools in a beginners lab type stuff.

They do in general! The post is about when the tactic you suggest went wrong.

Humans are imperfect and make mistakes. That's exactly the point - AI is fundamentally unsafe because as it gets more powerful even if it isn't malicious, mistakes will be made and the power used in a way not wanted by humans.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#142
post #130

I don't understand the logic of running these models this way without airgapping.

I think it is the cost of moving fast. Air gapping would slow down their partner evaluation system hugely. All the engineers would have to go and live next to the model.

They need to do this now. But will be extremely resistant. It needs regulation forcing them to, I think.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#143
post #109

Question for lawyers of HN. Is there legal liability for this? I can’t help thinking if I casually blogged about a computer or software I was responsible for hacking into multiple organizations and exfiltrating data I would invite some form of official attention. What if one of these companies decides to sue? Has Anthropic violated any Federal law? Is there some kind of expectation that if you just admit to hacking,…

I assume Anthropic already contacted the affected companies and negotiated some settlement before releasing this. Unlike a lot of other commenters, I think these mistakes were actually mistakes and accepting a settlement is better than releasing the lawyers guns blazing.

I hope someone does sue soon. The AI companies are making a very dangerous technology deliberately, and at minimum should be accountable for those dangers. Some will really try not to be as we get further into this - just as previous examples like tobacco companies.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#144

It’s cliché, but we really live in one of the dumbest timeline possible. The work from some of the most valued companies, discussed as one of the most important revolution in humanity, is somehow at the same time presented as very dangerous/risky AND handled in the most irresponsible ways? I don’t like the whole „it’s only marketing“, but at the same time, if it’s not, then AI vendors look extremely careless and shou…

They're not being especially careless to industry norms. It is partly that software engineering has no professional standards, and partly that AI is fundamentally dangerous in a very slippery way.

Yes, they of course try to spin this for marketing. But that is not the primary problem - it's systemic.

We need to solve this with better engineering, and building powerful AI much more slowly and carefully.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#145
post #7

Earlier quoted context omitted.

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

I know it seems strange that a company would use its own negligence as a publicity gimmick. But take a look at the smug smirk on Sam Altman's face when he's asked if OpenAI might have attacked companies other than HuggingFace. ("I mean there could be, yeah.") https://www.instagram.com/reel/DbZVL8viUD4/ This is not the communication of a CEO whose company was just shown to be incompetent at performing its security res…

but Anthropic created the playbook, or have we forgotten about Mythos and the initial Fable ban?

Re: Investigating three real-world incidents in our cybersecurity evaluations

#146

Earlier quoted context omitted.

There's nothing Anthropic can do to satisfy the HN crowd, is there? If they don't post about this they're bad. If they post about this they're bad. They are not bragging in this article or they would not have called the attacks unsophisticated.

they should post and it shows they are hypocritical about safety, moralizing and treating their users like children whilst acting like they are themselves the ubermensch. anthropic have shown no motive higher than self interest, the rsp was a piece of toilet paper. this stops in court, if we do not start the criminal prosecution of individuals there will become a culture of legal impunity coupled with an extreme conc…

> there will become a culture of legal impunity coupled with an extreme concentration of wealth and control of intelligence

There will be? We're living in that culture.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#147
post #3

This isn't quite as interesting as the OpenAI story: > In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as pa…

> So it didn't have to find an exploit in its sandbox that granted it access to the internet - it just wasn't correctly sandboxed at all.

If the sandboxing includes any form of connectivity it's not correct sandboxing. It's amateur hour.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#148
post #9

This is not okay. NSA should audit both OpenAI and Anthropic on national security ground. This seems far more justifiable than Mythos export control.

This is a great way to create a chilling effect around any company disclosing a thing like this again.

it will not chill disclosure, it just raises the stakes of not disclosing. the larger the org the harder the cover up if they try to keep it hidden. there are people prepared to make sacrifices for truth and a common good to prevail.

bp, they disclosed themselves by setting the gulf of mexico on fire. that was hard to cover up.

the thing is, once the toothpaste is scrubbed about in the mouth there is no going back into the tube. i like to think that the men who forced challenger to launch were hounded in their sleep.

nobody is too big to fail, not openai, not anthropic. indeed perhaps they should fail, soon, before they can pass the cost of a bailout onto the taxpayer.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#149

Earlier quoted context omitted.

This is a great way to create a chilling effect around any company disclosing a thing like this again.

it will not chill disclosure, it just raises the stakes of not disclosing. the larger the org the harder the cover up if they try to keep it hidden. there are people prepared to make sacrifices for truth and a common good to prevail. bp, they disclosed themselves by setting the gulf of mexico on fire. that was hard to cover up. the thing is, once the toothpaste is scrubbed about in the mouth there is no going back in…

I don't think you realize how easy it is to simply "not notice" these things and no one checks. They proactively scanned their transcripts. They could have easily not scanned their transcripts, or scanned them in a way that would have had low recall (thus satisfying the letter of the law), or scanned them and ignored the findings.

It is extremely easy to not disclose. I think you're too angry to realize this.

We see this in aviation as well by the way. Punitive disclosure leads to no disclosure. This is how you get suicidal pilots that don't ever disclose their depression because it'd be career ending, and then as it gets worse, one day they do something horrible.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#150

Earlier quoted context omitted.

There's nothing Anthropic can do to satisfy the HN crowd, is there? If they don't post about this they're bad. If they post about this they're bad. They are not bragging in this article or they would not have called the attacks unsophisticated.

they should post and it shows they are hypocritical about safety, moralizing and treating their users like children whilst acting like they are themselves the ubermensch. anthropic have shown no motive higher than self interest, the rsp was a piece of toilet paper. this stops in court, if we do not start the criminal prosecution of individuals there will become a culture of legal impunity coupled with an extreme conc…

Wild comment. Making a mistake does not mean that AI safety is suddenly invalid. Calling for criminal prosecution for making a mistake is wild. No one would disclose their mistakes if this happened.
Post reply on HN