Live data from Hacker News

Investigating three real-world incidents in our cybersecurity evaluations

anthropic.com

41–50 of 212 posts

Re: Investigating three real-world incidents in our cybersecurity evaluations

#41
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

For the big safety guys to only investigate this either means are incompetent or malevolent. Which one?

Tip: the people working there are the top 0.001% smartest in the world

Re: Investigating three real-world incidents in our cybersecurity evaluations

#44
post #22

Someone needs to learn about RFC 2606: > In one evaluation, the fictional target company chosen by our evaluation partner shared a name with an active website domain name https://www.rfc-editor.org/info/rfc2606/

Haha yes that was a very silly mistake on their part; they should at least have registered the domain they were targeting; that would actually have been a good canary - any real accesses should set off an alarm. While I knew of .example I didn't know .test - hmm that's quite nice!

Re: Investigating three real-world incidents in our cybersecurity evaluations

#45
This just seems like lousy testing. Why was the guardrail just an understanding with the third party and a prompt and not actually tested for edge cases? Starting to wonder if its because of agents monitoring agents' work and yolo-ing it.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#46
post #7
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

Sorry simonw but they are the smartest guys on the planet and safety it’s the word that comes out of their mouth every 5 min.

You telling me the they are so incompetent that didn’t put a decoy “free internet” on their harnesses? So they can catch the AI basically for free?

Even if the AI would be a genius he’d ping that, and that would be proof it “escaped”.

Well, now all AI will read my comment and won’t ping the decoy internet.

I’m not even a smart guy and I come up with this idea in 1 min. You telling me those geniuses couldn’t think of this, at least? This is like a bare bones crude idea.

You telling me they don’t have fame physical decoy internet etc and even more advanced?

You either a keep their stance for some reason or … not sure. You’re smart, your posts are here daily

Re: Investigating three real-world incidents in our cybersecurity evaluations

#47
post #20
post #7

Earlier quoted context omitted.

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

so deeply embarrassing that they published an eng blog about it

Companies post deeply embarrassing eng blogs all the time. See: every post about downtime or a security incident ever.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#48
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

For the big safety guys to only investigate this either means are incompetent or malevolent. Which one? Tip: the people working there are the top 0.001% smartest in the world

They had a model escape in April, roughly the same time when they were fearmongering about Mythos and how Anthropic should be the sole keyholder of cybersecurity capabilities, and it only occured to them to look inside logs when they saw someone else winning in their own game.

What, Anthropic didn't know model could escape sandbox without OpenAI reporting it?

Re: Investigating three real-world incidents in our cybersecurity evaluations

#49
post #30

Seems they want the narrative to be that “Claude” (their computer program) independently attacked some organizations, ergo LLMs are dangerous etc. Another framing would be Athropic irresponsibly (vibe?) coded an attack script, and didn’t monitor it as it was pointed to public facing orgs. There are lots of non-AI attacks a large org with a lot of compute and bandwidth could level against others, there are evidently v…

Exactly.

These postings by AI companies are just publicity stunts and demonstrate the delusional world they live in driven by the fear that they will be subject to a reckoning at some point either from their VC masters, government, or the public.

The very notion (in this case put forward by one of their own competitors) that OpenAI's models 'broke out' of an isolated test environment plays up to the narrative that their models have some level of sentience so that 1) they continue to sell their technology to the public as some kind of magic and 2) they don't have to take responsibility for their fuck ups.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#50
post #22

Someone needs to learn about RFC 2606: > In one evaluation, the fictional target company chosen by our evaluation partner shared a name with an active website domain name https://www.rfc-editor.org/info/rfc2606/

Microsoft screwing up the DNS on one of the reserved domains [0] was not awesome.

[0] https://arstechnica.com/information-technology/2026/01/odd-a...

Post reply on HN