Live data from Hacker News

Investigating three real-world incidents in our cybersecurity evaluations

anthropic.com

31–40 of 212 posts

Re: Investigating three real-world incidents in our cybersecurity evaluations

#31
post #19
post #7

Earlier quoted context omitted.

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

> I don't interpret it like that at all. This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! This just helps their (Anthropic) argument into persuading the US government into taking action into limiting powerful closed or open-weight models from being released without going thro…

They also gave access to Mythos (the Mythos) to some companies, based on... vibes.

Who knows how these companies are using it. If Anthropic can't effectively contain their own models, can the partners?

While the rest of us get fallbacks and warnings, not even being able to defend against the attacks they themselves are causing.

Do we really have to re-learn all the industry's knowledge the hard way?

Re: Investigating three real-world incidents in our cybersecurity evaluations

#32
post #18
post #12

> In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Of course and conveniently timed. With all these hacks being announced this is totally not a plan to scare governments in sho…

> The question now is why now? Because OpenAI figured out and wrote about their mistake first, which caused Anthropic to review their logs, and that's how they found out they had a similar problem back in April.

So a near trillion-dollar company doesn't have the basics of continuous security monitoring and threat-detection systems to catch and report this incident as soon as it is detected?

It doesn't help them to report serious incidents like this and it should be as soon as possible. This reactive investigation makes as if they ignored and sat on this issue, until a similar story from another lab made headlines first.

Would we have known about this issue if the OpenAI / Huggingface incident never happened?

Re: Investigating three real-world incidents in our cybersecurity evaluations

#33
post #20
post #7

Earlier quoted context omitted.

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

so deeply embarrassing that they published an eng blog about it

If they quietly brushed this under the rug - especially given the PyPI malware that was involved - it would be a huge scandal.

Disclosure is the only ethical response to this.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#34
There was no rush for this disclosure on their side. And they publish at a point where they have not yet taken corrective actions:

    > Some of the solutions here may even be simple fixes;
They are still throwing ideas. Why have they not made those simple fixes yet before disclosing?

Re: Investigating three real-world incidents in our cybersecurity evaluations

#35
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

I'm cynical as well, but the logical thing for them to do after the OpenAI/HF incident was to look at their systems for similar activity.

If they hadn't published this and instead it leaked out in two months we'd be slamming them for that as well.

They're stuck between a rock and a hard place, although they kind of put the rock there.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#37
post #20

Earlier quoted context omitted.

so deeply embarrassing that they published an eng blog about it

Right? 100% this is them trying to make gold out of turds.

There's nothing Anthropic can do to satisfy the HN crowd, is there? If they don't post about this they're bad. If they post about this they're bad.

They are not bragging in this article or they would not have called the attacks unsophisticated.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#39
> During that window, the package was downloaded and run on 15 real systems. One of these systems was a scanner belonging to a real security company (separate from the fictional company introduced in the scenario) that routinely installs Python packages and scans them for malware. When that company’s scanner installed the package, Claude’s hidden code executed. We believe the company’s security scanner treated PyPI packages as safe to install, and as a result, Claude was able to exfiltrate the company’s credentials to a collection point it had set up. Claude then used these credentials to access further infrastructure from this company.

lol. Natural stupidity remains undefeated!

Re: Investigating three real-world incidents in our cybersecurity evaluations

#40
post #29

Anthropic: "Look at how dangerous our models are!" Anthropic next week: "Why did you ban our models Mr Trump Daddy?"

So they should have kept silent about this? Then you'd be whining about that as well if it came to light
Post reply on HN