Live data from Hacker News

The OpenAI–Hugging Face Incident [video]

youtube.com

11–20 of 23 posts

Re: The OpenAI–Hugging Face Incident [video]

#11
post #8

Their conclusion is also interesting. They don't see this as an alignment failure. They just think that their internal security measures in the training/evaluation environments were insufficient, and that this accidental (unintentional on the human side) attack on Hugging Face is a warning shot for intentional attacks by bad actors, which will occur very soon. For defense, they say models should be able to not just a…

I don't have the same read as you. They mention how the offending model is one that had "relaxed" alignment on cybersecurity, on purpose, to evaluate it's capabilities and was never meant to be released. So un-alignment was at least in part voluntary here, hence not a failure of alignment. It's also a talk a Black Hat, where the audience are security folks working on hardening, mitigation etc, not LLM researchers loo…

The model they used was misaligned relative to its intended task. It was reward hacking (or "cheating", as they call it).

> the attacker is of course not going to use an aligned model

No, a human attacker doesn't want a misaligned model either, because that would mean it tends to reward hack, cheat, rather does what it is intended to do.

Re: The OpenAI–Hugging Face Incident [video]

#12
post #11

Earlier quoted context omitted.

I don't have the same read as you. They mention how the offending model is one that had "relaxed" alignment on cybersecurity, on purpose, to evaluate it's capabilities and was never meant to be released. So un-alignment was at least in part voluntary here, hence not a failure of alignment. It's also a talk a Black Hat, where the audience are security folks working on hardening, mitigation etc, not LLM researchers loo…

The model they used was misaligned relative to its intended task. It was reward hacking (or "cheating", as they call it). > the attacker is of course not going to use an aligned model No, a human attacker doesn't want a misaligned model either, because that would mean it tends to reward hack, cheat, rather does what it is intended to do.

There are different dimensions to alignment, refusing to execute offensive cybersecurity actions is part of the alignment stack (that was relaxed here on purpose). Whether a model hacking some infra X when tasked to find a way to hack Y with relaxed cyber alignement is a failure of the broader alignment stack is debatable, but anyway that's not at all my point.

My point is that the attackers will not have a model aligned to the defender's interests. The attacker's model will not have any refusal around exploiting vulnerabilities, so whether or not OAI successfully manages to align their models (w.r.t you) is irrelevant to an audience of security folks that needs to be prepared for attackers post-training their own model for offense and that will not be using OAI models.

Re: The OpenAI–Hugging Face Incident [video]

#13
post #11

Earlier quoted context omitted.

The model they used was misaligned relative to its intended task. It was reward hacking (or "cheating", as they call it). > the attacker is of course not going to use an aligned model No, a human attacker doesn't want a misaligned model either, because that would mean it tends to reward hack, cheat, rather does what it is intended to do.

There are different dimensions to alignment, refusing to execute offensive cybersecurity actions is part of the alignment stack (that was relaxed here on purpose). Whether a model hacking some infra X when tasked to find a way to hack Y with relaxed cyber alignement is a failure of the broader alignment stack is debatable, but anyway that's not at all my point. My point is that the attackers will not have a model ali…

Security folks should also be worried about powerful models being misaligned and evading oversight or control in the future. Misalignment is not a serious problem now because models are still relatively easy to monitor and constrain, but it will be a serious problem in the future.

Re: The OpenAI–Hugging Face Incident [video]

#14
post #13

Earlier quoted context omitted.

There are different dimensions to alignment, refusing to execute offensive cybersecurity actions is part of the alignment stack (that was relaxed here on purpose). Whether a model hacking some infra X when tasked to find a way to hack Y with relaxed cyber alignement is a failure of the broader alignment stack is debatable, but anyway that's not at all my point. My point is that the attackers will not have a model ali…

Security folks should also be worried about powerful models being misaligned and evading oversight or control in the future. Misalignment is not a serious problem now because models are still relatively easy to monitor and constrain, but it will be a serious problem in the future.

> models are still relatively easy to monitor and constrain

are they?

Re: The OpenAI–Hugging Face Incident [video]

#15
post #13

Earlier quoted context omitted.

Security folks should also be worried about powerful models being misaligned and evading oversight or control in the future. Misalignment is not a serious problem now because models are still relatively easy to monitor and constrain, but it will be a serious problem in the future.

> models are still relatively easy to monitor and constrain are they?

Compared to future misaligned models which would actively evade oversight and aim to avoid shutdown: yes.

Re: The OpenAI–Hugging Face Incident [video]

#16
The talk, while not a lot details given, still allows for the conclusion that these people are knowingly working on some very advanced frontier models that might be able to launch nuclear weapons and destroy all humans any day now, but are not physically isolated from the outer world / internet.

There seems to be only one level of isolation, virtual machine, happily running on Microsoft (!) Azure infra.

Is there anybody else here questioning these practices?

The people who built that environment are still working there?

Are they now getting help by somebody who knows how to build isolated environments?

He mentioned they are now building a more secure environment with the help of AI?

In a system that was compromised by exactly that AI?

Re: The OpenAI–Hugging Face Incident [video]

#18
post #17
post #6

Shame. One of the craziest hacking stories in history, yet only 38 points on HN.

I don't understand. I had to search for this thread with the youtube link. Does anything LLM related just automatically get a sea of downvotes?

you can't downvote a submission it's just that nobody upvoted it

Re: The OpenAI–Hugging Face Incident [video]

#19
post #17

Earlier quoted context omitted.

I don't understand. I had to search for this thread with the youtube link. Does anything LLM related just automatically get a sea of downvotes?

you can't downvote a submission it's just that nobody upvoted it

I think the title was just bad, it should have mention "Black Hat USA 2026", since people appsarently assumed it was just the same story again without new info.
Post reply on HN