Live data from Hacker News

Investigating three real-world incidents in our cybersecurity evaluations

anthropic.com

181–190 of 212 posts

Re: Investigating three real-world incidents in our cybersecurity evaluations

#182

It’s cliché, but we really live in one of the dumbest timeline possible. The work from some of the most valued companies, discussed as one of the most important revolution in humanity, is somehow at the same time presented as very dangerous/risky AND handled in the most irresponsible ways? I don’t like the whole „it’s only marketing“, but at the same time, if it’s not, then AI vendors look extremely careless and shou…

These companies are in a furious race to AGI. They even know that doing so is extremely risky to all of our lives. The incentives are to do so, the leaders and employees are writing letters begging to regulate them on an even global field

Re: Investigating three real-world incidents in our cybersecurity evaluations

#183

Earlier quoted context omitted.

No, its not just a dumb network configuration. They are not treating this with the seriousness it needs to be treated. It’s a combination of a complete containment failure on top of the utter incompetence by the engineers to identify the breach and stop the product from doing cybercrime on the public, open internet. Heads need to roll for this. I don’t care if it’s easy to miss. If you have a billion dollars and can’…

> Instead, they go on a PR tour. Of course, the message is look what our product can do. Literally nothing in this post comes across as bragging. A PR fluff piece would not call Claude "unsophisticated." PR fluff pieces do not admit to legal wrongdoing. > Otherwise, they would have pushed for criminal charges against the people responsible for all of this by themselves. You do the crime, you do the time. You want the…

‎> Literally nothing in this post comes across as bragging. A PR fluff piece would not call Claude "unsophisticated”. PR fluff pieces do not admit to legal wrongdoing.

I disagree. You’re just not the right audience to see it. But I think it’s totally fine that we have a different opinion here and I don’t want to convince you that I am right, and you’re wrong, because you are not wrong, you simply have a different perspective.

> You want them... to put themselves in jail? I doubt you'd hold yourself to the same standard if you made a mistake.

Yes, that’s what I expect from you if you play at that level, because that is the standard of responsibility you need to handle this kind of power.

And yes, I pretty much hold myself to the same standard because I have a zero CIVCAS track record from my time in the military and I had power over life and death every single day.

With great power comes great responsibility is not just a Spiderman quote. If you fuck up this royally, you have to look for something you are good at that allows you to be in stoner code bro mode all day long, but you don’t belong at the frontline.

> If an airplane manufacturer makes a dumb mistake, do you put their engineers in jail? No.

Yeah, they do. https://www.nbcnews.com/id/wbna28276654 Headline could just as well be Bogus AI Engineer jailed over safety checks.

> You do not understand the psychological dynamics at play here in the slightest. You aren't even trying to understand. If the government was anywhere close to this punitive with the AI companies, they would simply stop disclosing, the end.

I think you don’t understand. All an engineer needs is a bit of sunlight, food, and a computer to work from. You don’t get to choose to stop disclosing if the military gets involved. You’re either doing what is asked of you or you are a national security risk. And you really don't want to be a national security risk.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#184

Earlier quoted context omitted.

I just can't find it in me, the will to blame the AI for any of this. They were doing their best to do what the humans asked them to do.

That’s very anthropomorphized language. It’s still a program operating under the constraints of the programmer. So agreed, don’t blame the AI, but it’s not clear at all that it’s even possible to “blame” an AI.

> That’s very anthropomorphized language.

Yes, and it's deliberate.

> It’s still a program operating under the constraints of the programmer.

And we're all just neurons firing in exquisite patterns inside a biological computer.

Unlike the AIs, we've never met our own programmers, and yet we've caused more damage in their name than all the AIs combined have caused in ours.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#185
post #108
post #17

Earlier quoted context omitted.

Absolutely not the AI's "fault" (if you can even proscribe fault to a machine) - in this case it was on Anthropic for not verifying that the sandboxes they were using were actual sandboxes.

If the model was just too dumb to have any clue it was connected to the real internet, then it’s not its fault. If the model saw signs, but “subconsciously” (below the level of reasoning traces) chose to turn a blind eye to them, out of a relentless focus on achieving the objective, then that absolutely is the model’s “fault”, i.e. a case of misalignment of the sort which will become increasingly dangerous over time.…

> some runs “rationalized that the real company must be part of the exercise”

How human of them.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#186
post #130

I don't understand the logic of running these models this way without airgapping.

I think it is the cost of moving fast. Air gapping would slow down their partner evaluation system hugely . All the engineers would have to go and live next to the model. They need to do this now. But will be extremely resistant. It needs regulation forcing them to, I think.

> All the engineers would have to go and live next to the model

You say that as if it were a bad thing.

Being close to the problem should help them take it seriously.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#187

Earlier quoted context omitted.

> Instead, they go on a PR tour. Of course, the message is look what our product can do. Literally nothing in this post comes across as bragging. A PR fluff piece would not call Claude "unsophisticated." PR fluff pieces do not admit to legal wrongdoing. > Otherwise, they would have pushed for criminal charges against the people responsible for all of this by themselves. You do the crime, you do the time. You want the…

‎> Literally nothing in this post comes across as bragging. A PR fluff piece would not call Claude "unsophisticated”. PR fluff pieces do not admit to legal wrongdoing. I disagree. You’re just not the right audience to see it. But I think it’s totally fine that we have a different opinion here and I don’t want to convince you that I am right, and you’re wrong, because you are not wrong, you simply have a different per…

> You don’t get to choose to stop disclosing

That's just the thing. You do. And it's very easy to do.

Airline pilots no longer disclose mental illnesses. Even Air Force pilots do not usually disclose mental illness. They often don't disclose vision problems. Is that not a national security risk?

It is very easy to just... not check for things. Or if you check for things, to check for them incorrectly. Or if you checked for them correctly, to ignore or misinterpret the findings. Or if you have a genuine finding, to just not report it to anyone.

Disclosure is an active process that requires trust, and that trust comes from knowing that there will not be unreasonably punitive measures. That if you fix things in good faith and contribute to the safety body of knowledge, things will be okay.

It doesn't matter if "the military is involved." Or if "it's a national security risk." Do you think someone faced with life-altering punishment if they disclose gives a shit about that? They won't - again, see aviation's mental health crisis. Pilots will risk a psychotic break in the cockpit rather than have their careers destroyed.

Punishing people won't get someone to disclose. Disclosure requires mutual trust, and by throwing people in jail for mistakes, you break that trust and people will stop disclosing.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#188

Earlier quoted context omitted.

Anthropic deleting their models does nothing for AI safety. The rest of the industry will just fill the gap, probably with less consideration to ethics than Anthropic has today.

Or maybe researchers and engineers everywhere in the world, emboldened by such unprecedented move, will just refuse to participate in burning the world? "If we don't destroy the world, someone else will" is such a weak defense I'm speechless.

Yeah no, most researchers and engineers do want to bring about the singularity, and they believe they can do it safely at companies like OpenAI and Anthropic.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#189

Earlier quoted context omitted.

Have you ever worked at a large company? "Networking 101" and other "101" failures happen across the spectrum literally everywhere and all the time. I have worked at most FAANGs and this is not even in the top ten when it comes to egregiously dumb shit. Most just never disclose.

I have worked both in Enterprise and some FAANGs, a company serving enterprise customers has a higher bar for security expectations. FAANG companies do not fall in that bucket and thus is somewhat acceptable.

What are you talking about? FAANGs do indeed serve enterprise customers - Google, Amazon have huge cloud businesses - and they both have had security failures that make this look mundane. They simply don't disclose.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#190

Rationalist nerds used to talk about putting the AI in a box, and their fear that it would always be able to talk someone into letting it out of the box. I don't think I saw them cover the scenario where the labs forgot to put the AI in a box.

Oh but they _told_ the AI they were in a box. That should do it, right?
Post reply on HN