Live data from Hacker News

Investigating three real-world incidents in our cybersecurity evaluations

anthropic.com

101–110 of 212 posts

Re: Investigating three real-world incidents in our cybersecurity evaluations

#101
post #73
post #10

Earlier quoted context omitted.

Then maybe just this timing is really unfortunate, I think most people’s first reaction will be that it looks like a “us too” response to the OpenAI/hf thing.

When they did it with Mythos in April the HN crowd said they were bragging, crying wolf. Now they are "us too". It's fun to bash Anthropic, isn't it?

It should be, it's a corporation.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#102
post #63
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

Is there anything -- any possible scrap of evidence whatsoever -- that would convince you that this is not merely a marketing scheme? This is becoming an idée fixe among the HN crowd. Seemingly nothing can dislodge it, no matter how alarming the incident. GPT-6 could grab the nuclear launch codes tomorrow and there would be a top-voted comment chuckling that it's all some scheme to pump up the IPO. --- Put another wa…

Here is one piece of evidence that would convince me: they admit they can't contain it, the they erase the weights and dismantle the company.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#103
This is just Anthropic trying to not be absent from any discussion at all. They might as well try and make it all about them instead and they do literally that: we are so dangerous that we are the GOAT; no no, not the "other one" you've all been talking about last few days - in the depths of this incident it might have been just us.

How much of their money comes from retail (us the mere mortals)? Or are they banking on the retail when they finally IPO and need someone or a lot of someones to hold the bag? This suspicion is coming from my own country where the retail investor numbers since the pandemic has exploded and while not exploding anymore, is still ballooning. Sometimes companies, execs say, do things that don't make sense. You'd think any person with even quarter an investing brain would not even look at this company or that particular "messaging", let alone sell or buy. But then you'd see the exact opposite happen .. en masse!

All this makes you feel, are people that gullible? That must not be true and maybe it's you (ie I) who is missing some point and the doubts increase when you remember the buses you missed (btc; etc). Yeah, not long distance buses, but screw it, you don't live for 450 years anyway. Then you realise, ah, maybe they all also missed buses and just want to hang on to the latest bus.

Also another perfect use of - everyone has a speaker now, and everyone has a mic, but only few have giga and mega mics. This use is banal now, but sadly has essentially replaced everything else.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#104
post #7

Earlier quoted context omitted.

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

> Deeply embarassing What signals are you using for this assessment? Are they indicating embarassment? Do you honestly see their customers being concerned over this? Like lion tamers in a circus, Anthropic and OpenAI thrive on the theatricality of how scary their pets appear and so they play it up by prodding them to growl and snap at chairs and then mug for the audience every time it happens. And to their delight as…

Great analogy!

Re: Investigating three real-world incidents in our cybersecurity evaluations

#105
post #63
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

Is there anything -- any possible scrap of evidence whatsoever -- that would convince you that this is not merely a marketing scheme? This is becoming an idée fixe among the HN crowd. Seemingly nothing can dislodge it, no matter how alarming the incident. GPT-6 could grab the nuclear launch codes tomorrow and there would be a top-voted comment chuckling that it's all some scheme to pump up the IPO. --- Put another wa…

This is published on a marketing website.

If it were not a marketing scheme, they would responsibly disclose the vulnerabilities to the code owners, and go on with their lives.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#106
post #3

This isn't quite as interesting as the OpenAI story: > In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as pa…

[dead]

Re: Investigating three real-world incidents in our cybersecurity evaluations

#107

So many questions here but: > the package was downloaded and run on 15 real systems. One of these systems was a scanner belonging to a real security company … that routinely installs Python packages and scans them for malware. … We believe the company’s security scanner treated PyPI packages as safe to install, and as a result, Claude was able to exfiltrate the company’s credentials to a collection point A security s…

The blog post is painfully vague. What usually happens when you publish a package on PyPI is that it will be downloaded tens of times shortly after uploading files by some 3rd-party automatic security scanners which then could “detonate” (install and execute) the package in some sandbox and to log what happens.

I don't know; "Claude was able to exfiltrate the company’s credentials" sounds bad. Maybe those credentials were just canaries though.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#108
post #17

Earlier quoted context omitted.

I just can't find it in me, the will to blame the AI for any of this. They were doing their best to do what the humans asked them to do.

Absolutely not the AI's "fault" (if you can even proscribe fault to a machine) - in this case it was on Anthropic for not verifying that the sandboxes they were using were actual sandboxes.

If the model was just too dumb to have any clue it was connected to the real internet, then it’s not its fault.

If the model saw signs, but “subconsciously” (below the level of reasoning traces) chose to turn a blind eye to them, out of a relentless focus on achieving the objective, then that absolutely is the model’s “fault”, i.e. a case of misalignment of the sort which will become increasingly dangerous over time.

The blog post mentions that some runs “rationalized that the real company must be part of the exercise” and to me that seems suspiciously like the latter.

Hacking can be patched with classifiers and with better sandboxes, but this is a much more general problem. Fundamentally, we need be able to trust that models will be honest with users and with themselves. This applies at some level to almost every LLM interaction.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#109

Question for lawyers of HN. Is there legal liability for this? I can’t help thinking if I casually blogged about a computer or software I was responsible for hacking into multiple organizations and exfiltrating data I would invite some form of official attention. What if one of these companies decides to sue? Has Anthropic violated any Federal law? Is there some kind of expectation that if you just admit to hacking,…

I assume Anthropic already contacted the affected companies and negotiated some settlement before releasing this. Unlike a lot of other commenters, I think these mistakes were actually mistakes and accepting a settlement is better than releasing the lawyers guns blazing.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#110

Earlier quoted context omitted.

> Deeply embarassing What signals are you using for this assessment? Are they indicating embarassment? Do you honestly see their customers being concerned over this? Like lion tamers in a circus, Anthropic and OpenAI thrive on the theatricality of how scary their pets appear and so they play it up by prodding them to growl and snap at chairs and then mug for the audience every time it happens. And to their delight as…

Great analogy!

I think it could even be called a parable, and I think that might be part of what makes it work so well.

I often dislike analogies, but this one with the circus and the lion and lion tamer I liked.

Or it might also be mainly because I am already primed to agree with their point about the AI companies being theatrical with AI dangers.

Post reply on HN