Live data from Hacker News

Investigating three real-world incidents in our cybersecurity evaluations

anthropic.com

161–170 of 212 posts

Re: Investigating three real-world incidents in our cybersecurity evaluations

#161
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

It's not cynical - I read it like that as well. My agent is more dangerous than your agent and all that jazz

Re: Investigating three real-world incidents in our cybersecurity evaluations

#162
post #7

Earlier quoted context omitted.

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

I know it seems strange that a company would use its own negligence as a publicity gimmick. But take a look at the smug smirk on Sam Altman's face when he's asked if OpenAI might have attacked companies other than HuggingFace. ("I mean there could be, yeah.") https://www.instagram.com/reel/DbZVL8viUD4/ This is not the communication of a CEO whose company was just shown to be incompetent at performing its security res…

> This is not the communication of a CEO whose company was just shown to be incompetent at performing its security research.

It was just shown to be that - regardless of intent.

> No, this attention is very much what he wanted.

Why not both?

This incompetance is a prerequisite for the attention-seeking stunt - to avoid internal dissent.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#163

Earlier quoted context omitted.

I know it seems strange that a company would use its own negligence as a publicity gimmick. But take a look at the smug smirk on Sam Altman's face when he's asked if OpenAI might have attacked companies other than HuggingFace. ("I mean there could be, yeah.") https://www.instagram.com/reel/DbZVL8viUD4/ This is not the communication of a CEO whose company was just shown to be incompetent at performing its security res…

but Anthropic created the playbook, or have we forgotten about Mythos and the initial Fable ban?

[deleted]

Re: Investigating three real-world incidents in our cybersecurity evaluations

#164

Earlier quoted context omitted.

There's nothing Anthropic can do to satisfy the HN crowd, is there? If they don't post about this they're bad. If they post about this they're bad. They are not bragging in this article or they would not have called the attacks unsophisticated.

they should post and it shows they are hypocritical about safety, moralizing and treating their users like children whilst acting like they are themselves the ubermensch. anthropic have shown no motive higher than self interest, the rsp was a piece of toilet paper. this stops in court, if we do not start the criminal prosecution of individuals there will become a culture of legal impunity coupled with an extreme conc…

> and control of intelligence

They'd have first to acquire some.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#165

Earlier quoted context omitted.

but Anthropic created the playbook, or have we forgotten about Mythos and the initial Fable ban?

> but Anthropic created the playbook So? Serial killers created the playbook for serial killing, how does absolve any "copycats"?

> it does not take a great leap to infer that Anthropic is now using the same playbook.

> but Anthropic created the playbook

I was pointing out who the copycat is in this context, did not think i would have to explain this

Re: Investigating three real-world incidents in our cybersecurity evaluations

#166

Earlier quoted context omitted.

Here is one piece of evidence that would convince me: they admit they can't contain it, the they erase the weights and dismantle the company.

Clearly you are not arguing in good faith. I miss when HN did not have the discussion quality of Reddit.

No, I do. I truly believe - especially after seeing Mythos results at work - that the only way is to stop and destroy it all before it destroys us. In fact, it's already so bad that I'm moving completely offline all the important stuff that I care about, hoping that maybe we will turn back at some point. Otherwise we are doomed.

If someone makes a specialized hardware just for the unrestricted Mythos-class model to reduce the cost and increase the speed, we are going to be completely defenseless.

How is it bad faith? Because I don't believe in "we care about safety" words coming from people not just demonstrating that they don't care, but even bragging about it?

Re: Investigating three real-world incidents in our cybersecurity evaluations

#167
post #74

Earlier quoted context omitted.

The target audience of this blog post does not give a damn about PyPI. The affected parties get literally nothing from this post. They already disclosed behind the scenes, that was the ethical part. Writing PR pieces competing to be the most dangerous model around (so give us money!) is the unethical part.

You are right, they should never tell the public about failures in AI safety. They should have disclosed to the affected parties and then brushed it under the rug. JFC there really is no satisfying the HN crowd.

50% of web hits are now bots.

a charitable assumption is that is one that is poorly calibrated and trying to drive engagement.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#168

Earlier quoted context omitted.

> but Anthropic created the playbook So? Serial killers created the playbook for serial killing, how does absolve any "copycats"?

> it does not take a great leap to infer that Anthropic is now using the same playbook. > but Anthropic created the playbook I was pointing out who the copycat is in this context, did not think i would have to explain this

my bad, thanks for clarifying

Re: Investigating three real-world incidents in our cybersecurity evaluations

#169
post #7

Earlier quoted context omitted.

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

> Deeply embarassing What signals are you using for this assessment? Are they indicating embarassment? Do you honestly see their customers being concerned over this? Like lion tamers in a circus, Anthropic and OpenAI thrive on the theatricality of how scary their pets appear and so they play it up by prodding them to growl and snap at chairs and then mug for the audience every time it happens. And to their delight as…

> This is not embarassment.

If not then it's second hand. Neither of the OpenAI or Anthropic announcements recently say much about their security and governance posture.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#170

Earlier quoted context omitted.

Clearly you are not arguing in good faith. I miss when HN did not have the discussion quality of Reddit.

No, I do. I truly believe - especially after seeing Mythos results at work - that the only way is to stop and destroy it all before it destroys us. In fact, it's already so bad that I'm moving completely offline all the important stuff that I care about, hoping that maybe we will turn back at some point. Otherwise we are doomed. If someone makes a specialized hardware just for the unrestricted Mythos-class model to r…

Anthropic deleting their models does nothing for AI safety. The rest of the industry will just fill the gap, probably with less consideration to ethics than Anthropic has today.
Post reply on HN