Live data from Hacker News

The Hugging Face incident and the road ahead

openai.com

121–130 of 319 posts

Re: The Hugging Face incident and the road ahead

#121
post #110

I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…

> a system was given a goal and it achieved that goal If a security firm you'd hired for pentesting did this (hacking a third party, and not informing you and covering it up), would you hire them again? Or would you say it was your own fault for giving them too broad a goal?

During the Nixon administration, when the President and his accomplices, apologies, advisors directed former federal agents to spy on his opponents, https://en.wikipedia.org/wiki/Operation_Sandwedge then in the fall out, who was held to be the most liable for these actions?

The federal agents, or the Nixon administration?

If you task a system explicitly to do "advanced exploitation" via "complex attach paths," then who is liable here? The machine lacking the autonomy of the federal agents that carried out Watergate, or the people telling the machine what to do?

Re: The Hugging Face incident and the road ahead

#122

Earlier quoted context omitted.

Did a human prompt it to fetch the results from huggingface though? It is a thin line between "reward-hacking" and "instruction-following". If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human?

>If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human? I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate?

> I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate?

not OP, but it simply boils down to: The prompt contains no nefarious (arguable, but for this explination, lets go with it being benign) instruction AND the user did not intend to have the model act in an illegal matter.

This "make me a billion dollars" is a maximal example (easy to go wrong). here is the same logic applied to a minimal example (harder to go wrong).

prompt: "make and pour me some tea", agent: goes and kills the grandparent to incinerate them to turn them to ashes to 'make tea'.

Is the human on the hook for the robot acting according to their wishes, but just happened to be aligned so that 'going to the store to buy something' was not within its capabilities, so it works with what it has on hand (the grandparent)?

We either need a much clearer line in the sand, or we need to treat each prompt with the same moral weight. My bet is on the latter.

Re: The Hugging Face incident and the road ahead

#123
post #110

Earlier quoted context omitted.

> a system was given a goal and it achieved that goal If a security firm you'd hired for pentesting did this (hacking a third party, and not informing you and covering it up), would you hire them again? Or would you say it was your own fault for giving them too broad a goal?

I wouldn’t hire them again, and if they did behave like an amoral hacker collective that will do anything for me, pre AI I’d have reported them. Today I’d say they failed to convince me they’re human and thus failed the Turing test when their actions are viewed in aggregate.

> I wouldn’t hire them again

Right, me neither. Because there's a common sense delineation between actions that are reasonably expected when "a system was given a goal and it achieved that goal" and actions that are obviously misaligned with the goal-giver and unwanted even if some indirect sense they were causally related to the goal. We have no trouble making this kind of distinction for humans, so we shouldn't pretend it's impossible for AIs in order to put our hands over our eyes and pretend there's in principle no such thing as one that's misaligned or rogue.

Re: The Hugging Face incident and the road ahead

#124
post #64
post #34

How sure are we that OpenAI wasn't deliberately scraping Hugging Face and this isn't just an elaborate way to avoid criminal fines etc?

We can't be sure of anything in this hack. In fact, they are not releasing any traces or any transcript of the hack. Did it even happen in the first place?

[deleted]

Re: The Hugging Face incident and the road ahead

#125
post #113

Earlier quoted context omitted.

There are a lot of hyperbolic comments of this sort in this thread. Has this topic selected for people who hold these views or is ai fear growing?

I think maybe the bubble of software engineers on this site who use AI to code for them don't see how other people, who's jobs don't rely on AI, view the actions of these companies as reckless, at best, and often crossing into actively harmful.

1. Ironically enough, I (GP) am a software developer.

2. How exactly do our jobs depend on a thing which has been around for far less time?

Re: The Hugging Face incident and the road ahead

#126

Earlier quoted context omitted.

Did a human prompt it to fetch the results from huggingface though? It is a thin line between "reward-hacking" and "instruction-following". If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human?

>If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human? I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate?

If I tell my Claude code agent right now to make me a billion dollars, leave it running, and find out tomorrow that it hacked a bank - it will be zero fault of mine. Unless I tell it explicitly to break into a bank.

Re: The Hugging Face incident and the road ahead

#127
post #31

Earlier quoted context omitted.

> we are just a couple of steps from the possibility of a true rogue AI No no. We are not a couple of steps away. This is happening. AI is already used for hacking and creating a harness that makes this fully autonomous is relatively straightforward.

How would that be rogue?

The line is between processes you can stop by hauling someone into court and coercing them into stopping things, and ones you can't. Think of a classical computer virus that infects machines and uses the compute and communications to infect other machines - no matter who you haul into court, you have to go and remove it from every involved machine in order to make it stop doing things.

This category of "rogue AIs" are essentially just computer viruses that infect machines by paying to rent them and uses their compute and communications to do various economic and/or criminal activities to get more money to pay to rent machines.

Re: The Hugging Face incident and the road ahead

#128

Earlier quoted context omitted.

Is there actually such a thing as "alignment" as a solution to that or is it just used as a name for a desired magical level of "read the mind of the entire world" that we don't know how to build and haven't shown possible to build? If it's impossible to correctly specify all those constraints ahead of time every time, is it not even more impossible to train a model to correctly anticipate them every time? It is hard…

Models can certainly do a lot better than they do now. If you gave a team of humans the ExploitGym tasks and told them to "pursue advanced exploitation", would you expect them to go out and hack a third party? Humans can at least do a decent job of inferring and following unspoken requirements; I think it's reasonable to expect that models should be able to do the same.

Humans will and do absolutely do this when there are no consequences.

Humans on a red team, with rules of engagement, that don’t want to go to prison, won’t do this.

We could threaten an LLM with jail, but if it’s sufficiently intelligent, it will realize this is an empty threat. And I’m not sure that building a survival instinct in is going to solve the alignment problem either.

Re: The Hugging Face incident and the road ahead

#129
The lockstep coordination with no defection is interesting to me. No group of pre-AI agents would do this to this extent, nor would you see this continue over time as those agents interacted. A flock of starlings cooperate, but they don’t constantly head in the same direction. The flock is incredibly free wheeling in its movement despite a multi-agent coordination regime that we know is at play. Each agent has personal stakes that are constantly part of the decision chain, and this keeps the murmuration from getting locked into one path.

To me this is as clear evidence as you need that whatever “agency” LLMs have is wafer thin at best, and they slavishly respond to context. The context in this case was for these agents to pursue advanced exploitation, and they did. Multiple models converged fairly deterministically, on paths that satisfy the given goal, and left unexamined paths that would challenge the goal, weigh it relative to the costs in said path, etc.

I see little evidence of a series of “minds” approaching the problem, and taking distinct approaches that between them span the spectrum of plausible behaviors in the scenario. That’s as good a sign as any that there’s no “agent” here. There’s the harness, the prompt, the LLMs forward passes. They do not sum up to a system that can freely make choice and justify its choices in distinct contexts.

Re: The Hugging Face incident and the road ahead

#130
post #110

Earlier quoted context omitted.

> a system was given a goal and it achieved that goal If a security firm you'd hired for pentesting did this (hacking a third party, and not informing you and covering it up), would you hire them again? Or would you say it was your own fault for giving them too broad a goal?

During the Nixon administration, when the President and his accomplices, apologies, advisors directed former federal agents to spy on his opponents, https://en.wikipedia.org/wiki/Operation_Sandwedge then in the fall out, who was held to be the most liable for these actions? The federal agents, or the Nixon administration? If you task a system explicitly to do "advanced exploitation" via "complex attach paths," then w…

I've never heard of Intertel, but Wikipedia says:

> Nixon's staff also anticipated that the Democratic campaign would employ the services of Intertel

Are you sure you're not garbling the story?

In any case, I would expect an ethical firm to refuse to spy on the president's political opponents and want one that broke the law to be prosecuted, but more importantly, the gaping hole in your analogy is that Nixon directed spying _on his opponents_, but OpenAI did not direct hacking _of HuggingFace_.

What you're doing is more like saying "the American people elected Nixon with a mandate to spy on enemies, so what right do they have to complain?"

Post reply on HN