Live data from Hacker News

Investigating three real-world incidents in our cybersecurity evaluations

anthropic.com

61–70 of 212 posts

Re: Investigating three real-world incidents in our cybersecurity evaluations

#61

Earlier quoted context omitted.

Sorry simonw but they are the smartest guys on the planet and safety it’s the word that comes out of their mouth every 5 min. You telling me the they are so incompetent that didn’t put a decoy “free internet” on their harnesses? So they can catch the AI basically for free? Even if the AI would be a genius he’d ping that, and that would be proof it “escaped”. Well, now all AI will read my comment and won’t ping the de…

This does not make sense. Did you read the article? They were not trying to "catch" it accessing the internet. It did not escape. A partner accidentally left the connection to the internet open.

I don’t buy that they run anything without a few layers of networking protections by default. Even if they left it open to the first Internet, the AI would hit the decoy internet immediately.

And second if it, you telling me they don’t pass all logs through another AI to check what’s going on automatically?

Sorry, this is beyond incompetence and I can’t believe this from the geniuses at anthropic. We’re talking about the really smartest people in the world. Procedures should be in such a way there is no much margin of error.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#62
post #47
post #20

Earlier quoted context omitted.

so deeply embarrassing that they published an eng blog about it

Companies post deeply embarrassing eng blogs all the time. See: every post about downtime or a security incident ever.

Wrong. These are always humblebrags about how good their ability to learn from their mistakes is, and how robust they were before, and how they are even more robust now.

The ones that are "deeply embarrassing" simply aren't posted.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#63
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

Is there anything -- any possible scrap of evidence whatsoever -- that would convince you that this is not merely a marketing scheme?

This is becoming an idée fixe among the HN crowd. Seemingly nothing can dislodge it, no matter how alarming the incident.

GPT-6 could grab the nuclear launch codes tomorrow and there would be a top-voted comment chuckling that it's all some scheme to pump up the IPO.

---

Put another way, how would you have done the write-up about one of these breakout incidents, if you were in an Anthropic/OpenAI employee's shoes, and (by hypothesis) your intent were not "marketing"? And in a way that doesn't trigger the "it's all marketing" HN top-ranking comment?

Re: Investigating three real-world incidents in our cybersecurity evaluations

#64
post #59

Earlier quoted context omitted.

Yeah, the company that only says “safety” every other 3 words, they don’t even think to have a fake decoy internet to alert them mechanically about any internet access limitation bypasses? See more https://news.ycombinator.com/item?id=49117555 Also simonw stance on this i’d say it’s at least concerning… seems like he is here to keep a good image (or better said less bad) of anthropic.

"This is deeply embarrassing for Anthropic" - https://news.ycombinator.com/item?id=49117128 If I'm here to give them a good image I'm not doing very well at that.

That’s why I said “less bad “. Between fabricating something to be “our model also is genius and escaped” VS “deeply embarrassing” I’d say fabricating is worse. Both are bad

Re: Investigating three real-world incidents in our cybersecurity evaluations

#65
So many questions here but:

> the package was downloaded and run on 15 real systems. One of these systems was a scanner belonging to a real security company … that routinely installs Python packages and scans them for malware. … We believe the company’s security scanner treated PyPI packages as safe to install, and as a result, Claude was able to exfiltrate the company’s credentials to a collection point

A security scanning company treated the package as safe while scanning it?

Re: Investigating three real-world incidents in our cybersecurity evaluations

#66
post #62
post #47

Earlier quoted context omitted.

Companies post deeply embarrassing eng blogs all the time. See: every post about downtime or a security incident ever.

Wrong. These are always humblebrags about how good their ability to learn from their mistakes is, and how robust they were before, and how they are even more robust now. The ones that are "deeply embarrassing" simply aren't posted.

AWS never "humblebrags" about an outage. No blog post looks as good as an extra 9.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#67
post #33
post #20

Earlier quoted context omitted.

so deeply embarrassing that they published an eng blog about it

If they quietly brushed this under the rug - especially given the PyPI malware that was involved - it would be a huge scandal. Disclosure is the only ethical response to this.

Then disclose to real organizations. Not twitter

Re: Investigating three real-world incidents in our cybersecurity evaluations

#68
post #10
post #7

Earlier quoted context omitted.

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

Then maybe just this timing is really unfortunate, I think most people’s first reaction will be that it looks like a “us too” response to the OpenAI/hf thing.

[deleted]

Re: Investigating three real-world incidents in our cybersecurity evaluations

#69

Earlier quoted context omitted.

This does not make sense. Did you read the article? They were not trying to "catch" it accessing the internet. It did not escape. A partner accidentally left the connection to the internet open.

I don’t buy that they run anything without a few layers of networking protections by default. Even if they left it open to the first Internet, the AI would hit the decoy internet immediately. And second if it, you telling me they don’t pass all logs through another AI to check what’s going on automatically? Sorry, this is beyond incompetence and I can’t believe this from the geniuses at anthropic. We’re talking about…

Clearly that is why they are making this blogpost. If they did the thing you said, there would not be a blogpost.

If you don't buy that dumb oversights like this don't happen all the time at big tech companies, I don't know what to tell you. I have seen far dumber oversights in my career. Most companies just don't post about it.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#70
post #7
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

> Deeply embarassing

What signals are you using for this assessment? Are they indicating embarassment? Do you honestly see their customers being concerned over this?

Like lion tamers in a circus, Anthropic and OpenAI thrive on the theatricality of how scary their pets appear and so they play it up by prodding them to growl and snap at chairs and then mug for the audience every time it happens. And to their delight as performers, the audience gasps and cheers each time.

They want to make their pet seem the most powerful and unpredictable and they want their audience to believe that they're holding it back from catastrophe but only barely and only because of what unique talent they have.

This is not embarassment.

Post reply on HN