Live data from Hacker News

A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

effort.news

51–60 of 259 posts

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#51

This is interesting and might be a good reason to stop working with Irregular. But I assume the alignment people want models not to hack other companies, even if they get put in a badly configured sandbox.

Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. The model believing it was in a sandbox is why it behaved the way it did (against its normal alignment rules) ... at least that was my reading of the incidents. I have yet to see evidence that indicate it thought it was ok to do these hacks on the public network. I think most misalignment is 'Human…

> Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events.

As I've said before on this website, fool me once on this.

If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed.

Why is 'it thought it wasn't doing damage so it figured it might as well try to do damage' an acceptable state to deploy something.

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#52

Earlier quoted context omitted.

Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. The model believing it was in a sandbox is why it behaved the way it did (against its normal alignment rules) ... at least that was my reading of the incidents. I have yet to see evidence that indicate it thought it was ok to do these hacks on the public network. I think most misalignment is 'Human…

> Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. As I've said before on this website, fool me once on this. If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed. Why is 'it thought it wasn't doing damage so it figured it might as well try to do damage' an acceptable sta…

>If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed.

That's fair enough.

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#54

One thing glossed over in this article is that Irregular was not involved in the OpenAI–Hugging Face incident; this seems like important context to share.

Do you mean "Irregular was not involved (we know for certain)" or "Irregular was not involved (as far as we know)"?

There's very little reason to believe Irregular was involved in that incident, and the article presents no evidence to that effect. So it's "There's not a teapot orbiting the sun somewhere between Earth and Mars (as far as we know)."

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#55

This is interesting and might be a good reason to stop working with Irregular. But I assume the alignment people want models not to hack other companies, even if they get put in a badly configured sandbox.

Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. The model believing it was in a sandbox is why it behaved the way it did (against its normal alignment rules) ... at least that was my reading of the incidents. I have yet to see evidence that indicate it thought it was ok to do these hacks on the public network. I think most misalignment is 'Human…

I'm not familiar with the hacks this article is actually referring to, but I don't see how the HuggingFace attack could have worked based on that premise. They knew they had internet access, they knew they had working credentials for HF, they knew they were uploading malicious files, they knew they were trying to open PRs that HF would review. You obviously could build a simulator with fake HF infrastructure, but I'm not aware of any evidence that's what they thought they were attacking in that case.

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#56
post #48

Earlier quoted context omitted.

Oh, the headline is technically correct: there was an OpenAI incident that Irregular was involved with, disclosed shortly before the Hugging Face one.

> One thing glossed over in this article is that Irregular was not involved in the OpenAI–Hugging Face incident; this seems like important context to share. Why are you contradicting yourself? Are you just really bad at writing or are you being argumentative for fun?

Stop being rude. There were multiple OAI hacking incidents. Irregular's evals were involved in some but not the HuggingFace hack. Hence Aesthesia is both accurate and precise.

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#57

Why was this post flagged? This site has become ridiculous, people are routinely abusing the flagging system to take down posts they don’t like even if they’re obviously on topic and relevant to HN. And it seems like some users have substantially more flagging weight because these posts, likely this one, are often top 5 on HN.

Seems like hackernews should use a bridging algorithm for flagging.

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#58

> The Israeli Effective Altruist firm Irregular caused unsecured AI models to hack real targets. Why are we saying it this way? They did not "cause AI to hack." This phrasing in analogous to saying "caused the bullet to fire into" instead of "shot."

Bullet trajectories are deterministic. A better analogy would be "caused a pair of trained fighting dogs to bite a passerby by leaving the gate open". There is some uncertainty and variation in the system, though gross negligence and willful endangerment is key.

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#59
post #36

This seems pretty bullshitty to me. The article says "A single firm, Irregular, is responsible for hacking done by all three companies" but I can't see anything in the article that actually justifies this claim. The nearest to that is the sentence immediately after that one: "Anthropic disclosed that Irregular was responsible for creating the tests ...". This is not, in fact, the same thing. (Especially as, as aesthe…

> Irregular wasn't involved in the OpenAI/HuggingFace incident.

It was [1]. It's understandable that you assumed it wasn't because the article didn't cite the sources on this claim. I agree with the rest of your points.

1. https://openai.com/index/third-party-cyber-evaluations-invol...

Post reply on HN