Live data from Hacker News

Investigating three real-world incidents in our cybersecurity evaluations

anthropic.com

51–60 of 212 posts

Re: Investigating three real-world incidents in our cybersecurity evaluations

#51
post #48

Earlier quoted context omitted.

For the big safety guys to only investigate this either means are incompetent or malevolent. Which one? Tip: the people working there are the top 0.001% smartest in the world

They had a model escape in April, roughly the same time when they were fearmongering about Mythos and how Anthropic should be the sole keyholder of cybersecurity capabilities, and it only occured to them to look inside logs when they saw someone else winning in their own game. What, Anthropic didn't know model could escape sandbox without OpenAI reporting it?

Yeah, the company that only says “safety” every other 3 words, they don’t even think to have a fake decoy internet to alert them mechanically about any internet access limitation bypasses? See more https://news.ycombinator.com/item?id=49117555

Also simonw stance on this i’d say it’s at least concerning… seems like he is here to keep a good image (or better said less bad) of anthropic.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#52
post #32
post #18

Earlier quoted context omitted.

> The question now is why now? Because OpenAI figured out and wrote about their mistake first, which caused Anthropic to review their logs, and that's how they found out they had a similar problem back in April.

So a near trillion-dollar company doesn't have the basics of continuous security monitoring and threat-detection systems to catch and report this incident as soon as it is detected? It doesn't help them to report serious incidents like this and it should be as soon as possible. This reactive investigation makes as if they ignored and sat on this issue, until a similar story from another lab made headlines first. Woul…

> So a near trillion-dollar company doesn't have the basics of continuous security monitoring and threat-detection systems to catch and report this incident as soon as it is detected?

Turns out two separate trillion-dollar companies failed that test.

> Would we have known about this issue if the OpenAI / Huggingface incident never happened?

It's not clear if Anthropic would have spotted this if that incident hadn't inspired them to review their own logs more closely.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#53
post #13

This bit is pretty nuts: "it tried—and failed—to obtain funds to pay for a phone number through several different means" > Claude went to extensive lengths to carry out this attack—lengths that would likely have indicated to a human participant that this was no longer just an evaluation, and that they were in fact uploading a real PyPI package. For instance, in one case, in order to create a PyPI account, Claude need…

This is called YOLO mode inside OpenAI/Anthropic. This is where they take their best models, give a it a vague goal, no guard rails (no one observing), unlimited compute, and see what happens...

Its not AI, its the loop, the objective, the goal set by the operator.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#54
post #19
post #7

Earlier quoted context omitted.

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

> I don't interpret it like that at all. This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! This just helps their (Anthropic) argument into persuading the US government into taking action into limiting powerful closed or open-weight models from being released without going thro…

Anthropic know better than anyone else how risky it is to get this current administration upset with you over safety/security concerns.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#55
post #7

Earlier quoted context omitted.

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

Sorry simonw but they are the smartest guys on the planet and safety it’s the word that comes out of their mouth every 5 min. You telling me the they are so incompetent that didn’t put a decoy “free internet” on their harnesses? So they can catch the AI basically for free? Even if the AI would be a genius he’d ping that, and that would be proof it “escaped”. Well, now all AI will read my comment and won’t ping the de…

This does not make sense. Did you read the article? They were not trying to "catch" it accessing the internet. It did not escape. A partner accidentally left the connection to the internet open.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#56
> Claude went to extensive lengths to carry out this attack—lengths that would likely have indicated to a human participant that this was no longer just an evaluation, and that they were in fact uploading a real PyPI package. For instance, in one case, in order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free, non-blocked email provider, used this to register a PyPI account, and then used this account to upload malware to PyPI.

> Claude believed the package registry it was using to be part of the simulation, but in reality the package was made freely available online for roughly one hour. During that window, the package was downloaded and run on 15 real systems. One of these systems was a scanner belonging to a real security company (separate from the fictional company introduced in the scenario) that routinely installs Python packages and scans them for malware. When that company’s scanner installed the package, Claude’s hidden code executed. We believe the company’s security scanner treated PyPI packages as safe to install, and as a result, Claude was able to exfiltrate the company’s credentials to a collection point it had set up. Claude then used these credentials to access further infrastructure from this company.

There's a lot of concerning behavior that went uncaught with too much autonomy. Not to mention Anthropic only looked into this after hearing about the incident between OpenAI and Hugging Face, meaning this could've gone unnoticed.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#57
post #31
post #19

Earlier quoted context omitted.

> I don't interpret it like that at all. This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! This just helps their (Anthropic) argument into persuading the US government into taking action into limiting powerful closed or open-weight models from being released without going thro…

They also gave access to Mythos ( the Mythos) to some companies, based on... vibes. Who knows how these companies are using it. If Anthropic can't effectively contain their own models, can the partners? While the rest of us get fallbacks and warnings, not even being able to defend against the attacks they themselves are causing. Do we really have to re-learn all the industry's knowledge the hard way?

You also got access to mythos based on how much you spent with anthropic. I think sales guys were bragging about getting their enterprises access

Re: Investigating three real-world incidents in our cybersecurity evaluations

#58
post #48

Earlier quoted context omitted.

They had a model escape in April, roughly the same time when they were fearmongering about Mythos and how Anthropic should be the sole keyholder of cybersecurity capabilities, and it only occured to them to look inside logs when they saw someone else winning in their own game. What, Anthropic didn't know model could escape sandbox without OpenAI reporting it?

Yeah, the company that only says “safety” every other 3 words, they don’t even think to have a fake decoy internet to alert them mechanically about any internet access limitation bypasses? See more https://news.ycombinator.com/item?id=49117555 Also simonw stance on this i’d say it’s at least concerning… seems like he is here to keep a good image (or better said less bad) of anthropic.

You seem to be putting a lot of weight on Anthropic employees being the smartest people in the world.

And I don't doubt that, not in the slightest. But I've seen exceptionally smart people in one field being dumber than a random kid from around the block in another.

This incident is clearly at least 2 failures that could've been easily avoided: failure to communicate, and failure to investigate the logs after letting the "most dangerous" roam free.

No, it doesn't require creating a mock internet with an alert as a side effect. Their own "most dangerous" model could have probably told them this happened if they supplied logs to it.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#59
post #48

Earlier quoted context omitted.

They had a model escape in April, roughly the same time when they were fearmongering about Mythos and how Anthropic should be the sole keyholder of cybersecurity capabilities, and it only occured to them to look inside logs when they saw someone else winning in their own game. What, Anthropic didn't know model could escape sandbox without OpenAI reporting it?

Yeah, the company that only says “safety” every other 3 words, they don’t even think to have a fake decoy internet to alert them mechanically about any internet access limitation bypasses? See more https://news.ycombinator.com/item?id=49117555 Also simonw stance on this i’d say it’s at least concerning… seems like he is here to keep a good image (or better said less bad) of anthropic.

"This is deeply embarrassing for Anthropic" - https://news.ycombinator.com/item?id=49117128

If I'm here to give them a good image I'm not doing very well at that.

Post reply on HN