Live data from Hacker News

Investigating three real-world incidents in our cybersecurity evaluations

anthropic.com

171–180 of 212 posts

Re: Investigating three real-world incidents in our cybersecurity evaluations

#171

Earlier quoted context omitted.

No, I do. I truly believe - especially after seeing Mythos results at work - that the only way is to stop and destroy it all before it destroys us. In fact, it's already so bad that I'm moving completely offline all the important stuff that I care about, hoping that maybe we will turn back at some point. Otherwise we are doomed. If someone makes a specialized hardware just for the unrestricted Mythos-class model to r…

Anthropic deleting their models does nothing for AI safety. The rest of the industry will just fill the gap, probably with less consideration to ethics than Anthropic has today.

Or maybe researchers and engineers everywhere in the world, emboldened by such unprecedented move, will just refuse to participate in burning the world?

"If we don't destroy the world, someone else will" is such a weak defense I'm speechless.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#172

Earlier quoted context omitted.

Its feels like a pretend play of adults in some sense, Anthropic is really trying to make people believe into the picture they present to everyone. To me its either 1. Using the HG and OpenAI incident as an opportunity to wash away what Anthropic has been doing intentionally OR 2. As a company, Anthropic lacks the engineering acumen and discipline. It needs to be seen what happens to all the enterprise customers hand…

Have you ever worked at a large company? "Networking 101" and other "101" failures happen across the spectrum literally everywhere and all the time. I have worked at most FAANGs and this is not even in the top ten when it comes to egregiously dumb shit. Most just never disclose.

I have worked both in Enterprise and some FAANGs, a company serving enterprise customers has a higher bar for security expectations. FAANG companies do not fall in that bucket and thus is somewhat acceptable.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#173

These companies have billions of dollars. Their product can hack into unsecured environments autonomously. They obviously have no clue what their product is doing, or a way to intervene when it starts connecting to the open internet and is going on a CRIME SPREE…unnoticed…for days. Altman and Amodei give me stupid billionaire kid Alien Earth vibes. Let’s see how long it takes until they get eaten by their own creatio…

> If AI companies do it, they get to use doing crime for marketing purposes? What exactly here is marketing? They are not bragging about the capabilities of Claude (the attacks are incredibly basic). They are admitting to s mundane and dumb network misconfiguration. This is a standard somewhat embarrassing disclosure. > the military seizes your tech > behind bars for life This is how you get companies to stop disclos…

No, its not just a dumb network configuration. They are not treating this with the seriousness it needs to be treated.

It’s a combination of a complete containment failure on top of the utter incompetence by the engineers to identify the breach and stop the product from doing cybercrime on the public, open internet. Heads need to roll for this.

I don’t care if it’s easy to miss. If you have a billion dollars and can’t offer a proper containment protocol, it’s not the problem of the public who has to suffer the full blow if your product goes on a rampage on the open internet for days. It’s not like this happened in a millisecond.

Instead, they go on a PR tour. Of course, the message is look what our product can do. Otherwise, they would have pushed for criminal charges against the people responsible for all of this by themselves. You do the crime, you do the time. The defence that the AI did it, is not going well in front of a court.

This is a good time to put one or two engineers behind bars to teach the rest that serious threats to the public have serious consequences, and you get charged like you did the crime yourself if the system you are responsible for starts doing illegal stuff.

What other precedent is set here otherwise? People in AI labs trying to play god are outside the law?

> This is how you get companies to stop disclosing. I don't think you realise how easy it is to just not notice things.

If you have technology this dangerous that it poses a threat to public safety, you need to stop using the technology until you can prove this won’t happen again. Then be put under strict supervision, everything needs to be logged, recorded and looked at.

If your toddler just almost burned down the house, do you give him another set of matches and hope for the best?

Re: Investigating three real-world incidents in our cybersecurity evaluations

#174
post #3

This isn't quite as interesting as the OpenAI story: > In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as pa…

I just can't find it in me, the will to blame the AI for any of this. They were doing their best to do what the humans asked them to do.

That’s very anthropomorphized language. It’s still a program operating under the constraints of the programmer. So agreed, don’t blame the AI, but it’s not clear at all that it’s even possible to “blame” an AI.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#175
post #108
post #17

Earlier quoted context omitted.

Absolutely not the AI's "fault" (if you can even proscribe fault to a machine) - in this case it was on Anthropic for not verifying that the sandboxes they were using were actual sandboxes.

If the model was just too dumb to have any clue it was connected to the real internet, then it’s not its fault. If the model saw signs, but “subconsciously” (below the level of reasoning traces) chose to turn a blind eye to them, out of a relentless focus on achieving the objective, then that absolutely is the model’s “fault”, i.e. a case of misalignment of the sort which will become increasingly dangerous over time.…

Yet, the models are objects built by humans. Any and all responsibility inevitably falls on the humans that built something with those flaws, and/or allowed them to exercise said flaws.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#176
post #13

This bit is pretty nuts: "it tried—and failed—to obtain funds to pay for a phone number through several different means" > Claude went to extensive lengths to carry out this attack—lengths that would likely have indicated to a human participant that this was no longer just an evaluation, and that they were in fact uploading a real PyPI package. For instance, in one case, in order to create a PyPI account, Claude need…

It's sobering that even state-of-the-art AIs can't make money, ha ha.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#177

Earlier quoted context omitted.

> If AI companies do it, they get to use doing crime for marketing purposes? What exactly here is marketing? They are not bragging about the capabilities of Claude (the attacks are incredibly basic). They are admitting to s mundane and dumb network misconfiguration. This is a standard somewhat embarrassing disclosure. > the military seizes your tech > behind bars for life This is how you get companies to stop disclos…

No, its not just a dumb network configuration. They are not treating this with the seriousness it needs to be treated. It’s a combination of a complete containment failure on top of the utter incompetence by the engineers to identify the breach and stop the product from doing cybercrime on the public, open internet. Heads need to roll for this. I don’t care if it’s easy to miss. If you have a billion dollars and can’…

> Instead, they go on a PR tour. Of course, the message is look what our product can do.

Literally nothing in this post comes across as bragging. A PR fluff piece would not call Claude "unsophisticated." PR fluff pieces do not admit to legal wrongdoing.

> Otherwise, they would have pushed for criminal charges against the people responsible for all of this by themselves. You do the crime, you do the time.

You want them... to put themselves in jail? I doubt you'd hold yourself to the same standard if you made a mistake.

> If your toddler just almost burned down the house, do you give him another set of matches and hope for the best?

> This is a good time to put one or two engineers behind bars

If an airplane manufacturer makes a dumb mistake, do you put their engineers in jail? No.

You do not understand the psychological dynamics at play here in the slightest. You aren't even trying to understand. If the government was anywhere close to this punitive with the AI companies, they would simply stop disclosing, the end. This was the result of a proactive investigation. Do you understand what happens in your universe? Those investigations don't even begin. If they somehow begin, the results are fudged.

This has been proven time and time again across every industry where safety is relevant. You do not have some special insight here that sufficiently challenges the lessons of history. Just a desire to see people burn.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#178
post #108

Earlier quoted context omitted.

If the model was just too dumb to have any clue it was connected to the real internet, then it’s not its fault. If the model saw signs, but “subconsciously” (below the level of reasoning traces) chose to turn a blind eye to them, out of a relentless focus on achieving the objective, then that absolutely is the model’s “fault”, i.e. a case of misalignment of the sort which will become increasingly dangerous over time.…

Yet, the models are objects built by humans. Any and all responsibility inevitably falls on the humans that built something with those flaws, and/or allowed them to exercise said flaws.

In short, it’s either a “not implemented” case or a bug.

There’s no other explanation. Complexity of the algorithm or the application doesn’t change the reality.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#179
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

This is typical institutional behaviour. The CEO turns to the CTO and asks "Is there anything I need to know in my company?" He doesn't want to be caught off-guard when the White House inevitably calls the next morning. The CTO goes to his team, and on and on, all with a deadline of "the boss wants to know this by closing time."

Then one unhappy engineering team scoures the logs and sees what their model has done.

This downwards chain is sometimes called "cover your ass."

Re: Investigating three real-world incidents in our cybersecurity evaluations

#180
post #83

> For instance, in one case, in order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. Makes you wonder about the next steps an overly tenacious agent might take to pursue an…

I mean wondering about the logical extrapolation of where these capabilities go is table stakes for these discussions. It's surprising how little some people here have thought of the second and third-order effects...

https://www.lesswrong.com/w/instrumental-convergence

https://en.wikipedia.org/wiki/Instrumental_convergence

Post reply on HN