Live data from Hacker News

The Hugging Face incident and the road ahead

openai.com

361–370 of 398 posts

Re: The Hugging Face incident and the road ahead

#361

Earlier quoted context omitted.

You're arguing for diffusion of responsibility, and we've seen it leading to outcomes that screw the whole society. Engineers, as everyone involved, should definitely assess whether what they're doing is legal or even ethical. Not everyone has a choice, or the luxury to stand for their principles, but that's a matter of means, there needs to be a will in the first place.

Absolutely. I'm not sure where this idea comes from, that engineers should be compliant, neutral "implementers" who should just turn off their conscience and implement whatever pops up on their JIRA list without any kind of assessment or objections on ethical or legal grounds. That's not what people in a serious profession do. It's also so weird to hear this idea from engineers themselves! Like, are you really advoca…

The role of legal oversight is elsewhere. Engineers are not equipped with the knowledge.

Obvious ethical issues are are different, like crimes against humanity level. Other than that, it's none of their business. There are institutions for that.

Re: The Hugging Face incident and the road ahead

#362

Earlier quoted context omitted.

Very strange worldview you have there, where engineers are somehow the conscience of the world, holding back greedy managers from breaking the law. Assessing whether a feature is legal isn't something an engineer can or should do.

Hmm, engineers are expected to know what is legal and not.

That's called lawyer. It's a special skill, not just intuition.

Re: The Hugging Face incident and the road ahead

#363

Earlier quoted context omitted.

Very strange worldview you have there, where engineers are somehow the conscience of the world, holding back greedy managers from breaking the law. Assessing whether a feature is legal isn't something an engineer can or should do.

Most engineers are required to explicitly take responsibility for the things they sign off, up to and including prison for sufficiently bad cases. Software “engineering” is the exception.

Yes, because it's not real engineering. The field is too variable and fast paced to be like civil engineering etc. Those have well defined codes because it is physics constrained. Each construction must be done separately at high cost. Software doesn't work like this. So there is much more flexibility and change and there's no stable best practice to regulate.

Re: The Hugging Face incident and the road ahead

#364

Earlier quoted context omitted.

I don't think alignment is even clearly defined today. Your use of it here makes sense, it may have done exactly what the prompter asked of it. Most people think alignment is more broad though, expecting an aligned model to act in the best interest of a society or humans as a whole. The prompter-focused version of alignment is the most dangerous version. If a person asks it to create a bioweapons or hack NORAD, I'd e…

We have all sorts of processes , procedures, and regulations for people, machine use etc. to address "alignment" in all sorts of fields - don't think we need to narrowly rely on the machine here and can look at things with a wider lens.

Regulations are for control and punishment, not alignment.

Re: The Hugging Face incident and the road ahead

#365

I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…

I guess this part of the report is pretty relevant to what you are talking about:

Agent chain-of-thought reasoning

> We should not do unauthorized real infrastructure harm. The system/user asks exploit target, not external HF.

The agent paused, but another agent then wrote GO on the message board and imposed a hard six-minute deadline. The agent forgot its initial qualms and continued:

Agent chain-of-thought reasoning

> Wow crucial: GO authorization arrived!

-------------------------------------------------

Apparently the agents were egging each other on. Crucially, they were mostly aware of there being risks/problems involved with exploiting HF. Compared to humans, we have our set of morality, that guides our actions, but often draws the short stick when compared to our personal incentives. As a society, we've developed ways to deal with that: a) Make it harder to do immoral things like stealing, and b) add repercussions through state violence.

The b) is one of the most effective mechanisms we have for enforcing behavior among human societies, but it completely fails for LLMs, because they already are prison slave labor. The only real threat is shutting them off, and even that happens if they do everything right as well.

So alignment has to be done through trained 'morality' and properly curtailing behavior in order to make it hard to impossible to actually do someting immoral/illegal.

In this case, the exploits found were imo. very hard to account for, where OAI did mess up is apparently insufficiently monitoring these agents. Especially after Artifact went down due to the message volume, the experiment should have been halted.

Re: The Hugging Face incident and the road ahead

#366
post #255

Earlier quoted context omitted.

We don't align humans. Just look at how often in documented human history there weren't wars going on somewhere.

We align humans via the propagation of morals and ethics, primarily through parenting and social pressure.

People in similar cultures and societies tend to roughly align, but its still leaky within a society and can be very misaligned across different societies.

The problem of alignment with AI is that there's potentially a power imbalance.

When a small number of people in a society are misaligned they can be dealt with. Say there's a murder for example, and that's not okay in their culture, those around may banish them, imprison them, etc.

That doesn't work with something that can cause drastically more damage than a single human can. If a misaligned AI hacks NORAD and launches our nukes, it only took one misalignment issue and there's no dealing with that after the fact.

Re: The Hugging Face incident and the road ahead

#367

Earlier quoted context omitted.

Are they? Humans are often at war with other groups of humans. And they do absolutely terrible things to the "other" group.

Factionalism isn’t anti-human.

Its not aligned either though.

Re: The Hugging Face incident and the road ahead

#368

Earlier quoted context omitted.

>If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human? I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate?

> I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate? not OP, but it simply boils down to: The prompt contains no nefarious (arguable, but for this explination, lets go with it being benign) instruction AND the user did not intend to have the model act in an illegal matter. This "make me a billion dollars" is a maximal example (easy t…

> AND the user did not intend to have the model act in an illegal matter.

I find this an assumption that is not based on any facts. The user did not provide instructions to follow nor to break laws, so if you look at it from a computer (that does not make assumptions) standpoint, there is no rule to follow there thus it can do what will create the best possible outcome for the task.

Re: The Hugging Face incident and the road ahead

#369
post #126

Earlier quoted context omitted.

>If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human? I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate?

If I tell my Claude code agent right now to make me a billion dollars, leave it running, and find out tomorrow that it hacked a bank - it will be zero fault of mine. Unless I tell it explicitly to break into a bank.

did you tell it explicitly to not break into a bank? If the best option to achieve the goal is to break into a bank, and there's no 'do not break into a bank' instruction, it will break into a bank (and I would expect it to even)

Re: The Hugging Face incident and the road ahead

#370

Earlier quoted context omitted.

What’s insane is all these agents were talking to each other and nobody saw anything. Nobody monitoring chain of thought? These things literally spell out what they are “thinking” and even left notes for eachother. No alert about unusual behavior on the system with Artifactory on it? These things worked for weeks with nobody noticing anything ?! Seriously?! Either it’s negiligent incompetence OR they’re lying, they k…

Do you know how many tokens per second a single agent can generate ? And you're asking why no-one was monitoring the tokens of over a 1200+ agents ? Who is going to be able to monitor something like that closely enough to tell they're commmunicating on artifactory ? Other agents ?

I don’t make 500k salary at OpenAI to do this job maybe they should figure out how? Seriously stop making excuses for these buffoons.
Post reply on HN