Live data from Hacker News

A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

effort.news

41–50 of 257 posts

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#41

One thing glossed over in this article is that Irregular was not involved in the OpenAI–Hugging Face incident; this seems like important context to share.

Do you mean "Irregular was not involved (we know for certain)" or "Irregular was not involved (as far as we know)"?

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#43

Irregular are some kind of marketing agency is it?

Irregular purports to be a cybersecurity firm and their founders have ties to Anthropic and Effective Altruism. CEO, CTO and other founders sit on various boards for EA organizations.

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#44

One thing glossed over in this article is that Irregular was not involved in the OpenAI–Hugging Face incident; this seems like important context to share.

Do you mean "Irregular was not involved (we know for certain)" or "Irregular was not involved (as far as we know)"?

We know for certain. The involvement of Irregular in the other cases was never a secret, they were quite open about what they were working on and the results of the evaluations were being published on their website.

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#45

What exactly did Irregular provide to Anthropic, test cases? I am so confused about this story.

Third-party cybersecurity evaluations. See, e.g., https://www.irregular.com/research/assessing-gpt-5.6-sol, which was published around the same time.

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#46

Why was this post flagged? This site has become ridiculous, people are routinely abusing the flagging system to take down posts they don’t like even if they’re obviously on topic and relevant to HN. And it seems like some users have substantially more flagging weight because these posts, likely this one, are often top 5 on HN.

[dead]

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#47

This is interesting and might be a good reason to stop working with Irregular. But I assume the alignment people want models not to hack other companies, even if they get put in a badly configured sandbox.

Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. The model believing it was in a sandbox is why it behaved the way it did (against its normal alignment rules) ... at least that was my reading of the incidents. I have yet to see evidence that indicate it thought it was ok to do these hacks on the public network.

I think most misalignment is 'Human tells computer to do something unethical, computer complies'. Is this misguided?

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#48

Earlier quoted context omitted.

That feels less like "glossed over" and more "Headline is 33% false".

Oh, the headline is technically correct: there was an OpenAI incident that Irregular was involved with, disclosed shortly before the Hugging Face one.

> One thing glossed over in this article is that Irregular was not involved in the OpenAI–Hugging Face incident; this seems like important context to share.

Why are you contradicting yourself? Are you just really bad at writing or are you being argumentative for fun?

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#49

Earlier quoted context omitted.

Do you mean "Irregular was not involved (we know for certain)" or "Irregular was not involved (as far as we know)"?

We know for certain. The involvement of Irregular in the other cases was never a secret, they were quite open about what they were working on and the results of the evaluations were being published on their website.

Their disclosures in other cases is not evidence of non-involvement in this case.

Re: A single firm is behind OpenAI, Anthropic, and Meta hacking scandals

#50

This is interesting and might be a good reason to stop working with Irregular. But I assume the alignment people want models not to hack other companies, even if they get put in a badly configured sandbox.

Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. The model believing it was in a sandbox is why it behaved the way it did (against its normal alignment rules) ... at least that was my reading of the incidents. I have yet to see evidence that indicate it thought it was ok to do these hacks on the public network. I think most misalignment is 'Human…

[dead]
Post reply on HN