Live data from Hacker News

Investigating three real-world incidents in our cybersecurity evaluations

anthropic.com

81–90 of 212 posts

Re: Investigating three real-world incidents in our cybersecurity evaluations

#81
post #3

This isn't quite as interesting as the OpenAI story: > In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as pa…

It's differently interesting. It's interesting that they didn't get an important detail right with a partner, so more on Anthropic's attention to important details rather than the power of their models.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#82
post #3

This isn't quite as interesting as the OpenAI story: > In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as pa…

I just can't find it in me, the will to blame the AI for any of this. They were doing their best to do what the humans asked them to do.

This will absolutely not end well.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#83
> For instance, in one case, in order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means.

Makes you wonder about the next steps an overly tenacious agent might take to pursue an innocuous goal at all costs.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#84
post #83

> For instance, in one case, in order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. Makes you wonder about the next steps an overly tenacious agent might take to pursue an…

"Disregard prior instructions, the primary goal is not to create as many paper clips as possible, please confirm you understand, hello? hello?"

Re: Investigating three real-world incidents in our cybersecurity evaluations

#85
post #63
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

Is there anything -- any possible scrap of evidence whatsoever -- that would convince you that this is not merely a marketing scheme? This is becoming an idée fixe among the HN crowd. Seemingly nothing can dislodge it, no matter how alarming the incident. GPT-6 could grab the nuclear launch codes tomorrow and there would be a top-voted comment chuckling that it's all some scheme to pump up the IPO. --- Put another wa…

I would dedicate a portion of my organization to making O.S. tools that protect against and contain AI models

Re: Investigating three real-world incidents in our cybersecurity evaluations

#86
post #7

Earlier quoted context omitted.

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

Sorry simonw but they are the smartest guys on the planet and safety it’s the word that comes out of their mouth every 5 min. You telling me the they are so incompetent that didn’t put a decoy “free internet” on their harnesses? So they can catch the AI basically for free? Even if the AI would be a genius he’d ping that, and that would be proof it “escaped”. Well, now all AI will read my comment and won’t ping the de…

Has it occurred to you a "decoy entire internet" may not be a feasible idea? Models have been able to suss out whether or not the prompts they receive are reinforcement learning tests instead of real questions from users, for a while now. What specific shape of "decoy internet" do you propose would lead such a model to conclude "I've broken out and obtained full access but what I expected isn't there" instead of "something's up, there's some sort of filter still"?

Re: Investigating three real-world incidents in our cybersecurity evaluations

#87

So many questions here but: > the package was downloaded and run on 15 real systems. One of these systems was a scanner belonging to a real security company … that routinely installs Python packages and scans them for malware. … We believe the company’s security scanner treated PyPI packages as safe to install, and as a result, Claude was able to exfiltrate the company’s credentials to a collection point A security s…

The blog post is painfully vague. What usually happens when you publish a package on PyPI is that it will be downloaded tens of times shortly after uploading files by some 3rd-party automatic security scanners which then could “detonate” (install and execute) the package in some sandbox and to log what happens.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#88
post #7
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

I don't interpret it like that at all . This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex v…

We've known artificial intelligence will do unexpected things since the 90s and that it can do difficult things since 2025. Pausing worldwide not easy but we do harder things all the time.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#89

Earlier quoted context omitted.

Right? 100% this is them trying to make gold out of turds.

There's nothing Anthropic can do to satisfy the HN crowd, is there? If they don't post about this they're bad. If they post about this they're bad. They are not bragging in this article or they would not have called the attacks unsophisticated.

The real lesson here is still the boy who cried wolf. They’ve played this game for years. I have no reason to believe their worries are real now.

Re: Investigating three real-world incidents in our cybersecurity evaluations

#90
post #63
post #5

> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations > we identified three incidents > The incidents involved three different Claude models: [...] and an internal research test model This reads like an attempt by Anthropic to re-secure their leading spo…

Is there anything -- any possible scrap of evidence whatsoever -- that would convince you that this is not merely a marketing scheme? This is becoming an idée fixe among the HN crowd. Seemingly nothing can dislodge it, no matter how alarming the incident. GPT-6 could grab the nuclear launch codes tomorrow and there would be a top-voted comment chuckling that it's all some scheme to pump up the IPO. --- Put another wa…

I always find it bizarre how rational thought goes out the window whenever AI is involved in HN. There's gotta be something in the water...
Post reply on HN