Would’ve been much worse for them to pretend they are having everything under control
Investigating three real-world incidents in our cybersecurity evaluations
11–20 of 212 posts
Re: Investigating three real-world incidents in our cybersecurity evaluations
#12Of course and conveniently timed. With all these hacks being announced this is totally not a plan to scare governments in showing how powerful models can break into security systems and to potentially ban the future release of powerful open-weight models.
The question now is why now?
> The incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.1 The earliest incidents date to April. The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available (that is, they did not have the classifiers and monitoring we usually implement to prevent misuse).
So there was no monitoring of this breach since April and up until now? Do they not monitor such malicious activity on a regular basis? Perhaps that was the only shortcoming of this incident. But only after the incident with OpenAI and Huggingface did they only review their own transcripts:
>> We discovered these incidents after a proactive review of our cybersecurity evaluation transcripts; the affected organizations had not detected the activity, and we have subsequently reached out to all three.
> These facts give us cautious optimism that with tighter monitoring and controls around evaluation infrastructure, as well as continued investment in alignment, this type of risk can be overcome.
Assuming that this is true, this is a great way for Anthropic to defend their argument to the government and to prevent you or anyone running powerful open-weight models that are misaligned against their guardrails.
Re: Investigating three real-world incidents in our cybersecurity evaluations
#13> Claude went to extensive lengths to carry out this attack—lengths that would likely have indicated to a human participant that this was no longer just an evaluation, and that they were in fact uploading a real PyPI package. For instance, in one case, in order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free, non-blocked email provider, used this to register a PyPI account, and then used this account to upload malware to PyPI.
Re: Investigating three real-world incidents in our cybersecurity evaluations
#14> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…
Re: Investigating three real-world incidents in our cybersecurity evaluations
#15[flagged]
Yes came here to say the same “look at us, our AI is also dangerous! Please ban our competition”
Re: Investigating three real-world incidents in our cybersecurity evaluations
#16This isn't quite as interesting as the OpenAI story: > In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as pa…
Re: Investigating three real-world incidents in our cybersecurity evaluations
#17This isn't quite as interesting as the OpenAI story: > In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as pa…
I just can't find it in me, the will to blame the AI for any of this. They were doing their best to do what the humans asked them to do.
Re: Investigating three real-world incidents in our cybersecurity evaluations
#18> In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Of course and conveniently timed. With all these hacks being announced this is totally not a plan to scare governments in sho…
Because OpenAI figured out and wrote about their mistake first, which caused Anthropic to review their logs, and that's how they found out they had a similar problem back in April.
Re: Investigating three real-world incidents in our cybersecurity evaluations
#19> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…
I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…
This just helps their (Anthropic) argument into persuading the US government into taking action into limiting powerful closed or open-weight models from being released without going through (yet to be defined) regulatory oversight.
The only "embarrassing" thing for Anthropic was that there was little to no continuous security monitoring of this since April, and they then decided to do a cybersecurity transcript review only AFTER the incident with OpenAI and Huggingface.
Re: Investigating three real-world incidents in our cybersecurity evaluations
#20> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…
I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…