Live data from Hacker News

The Hugging Face incident and the road ahead

openai.com

421–429 of 429 posts

Re: The Hugging Face incident and the road ahead

#421

1. They TOLD the model to "pursue advanced exploitation" to quantify its "cyber capabilities" (whatever that means). 2. The model pursues advanced exploitation. 3. "There was a incident due to dangerous actions taken by the model that no human directed" This is basically the pre-cursor of the paperclip maximizer [0], the AI executes the given order to an extend that was not considered in the order, now suddenly no-on…

"pursuing advanced exploitation" when explicitly given a sandbox in a VM and a benchmark problem involving a cyber exploit very clearly excludes hacking third parties. I think writing out the event in a 3 point list like that is disingenuous.

This is basic alignment, not even a tricky or ambiguous case.

I do very much agree with your take on culpability/military parallels, though.

Re: The Hugging Face incident and the road ahead

#422

Earlier quoted context omitted.

> I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate? not OP, but it simply boils down to: The prompt contains no nefarious (arguable, but for this explination, lets go with it being benign) instruction AND the user did not intend to have the model act in an illegal matter. This "make me a billion dollars" is a maximal example (easy t…

> AND the user did not intend to have the model act in an illegal matter. I find this an assumption that is not based on any facts. The user did not provide instructions to follow nor to break laws, so if you look at it from a computer (that does not make assumptions) standpoint, there is no rule to follow there thus it can do what will create the best possible outcome for the task.

We presume innocence where I'm from. If a message contains nothing nefarious, nefarious intent can not be assumed.

Re: The Hugging Face incident and the road ahead

#423
post #253

Earlier quoted context omitted.

> I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate? not OP, but it simply boils down to: The prompt contains no nefarious (arguable, but for this explination, lets go with it being benign) instruction AND the user did not intend to have the model act in an illegal matter. This "make me a billion dollars" is a maximal example (easy t…

"But your honour, my horseless carriage was not designed to hit children!"

right! Which is why we punish the driver and not the manufacturer.

Re: The Hugging Face incident and the road ahead

#425
post #400

Earlier quoted context omitted.

>How were they supposed to know about "previously unknown vulnerabilities"? You don't. That's why you unplug the Ethernet cable.

Have you done that to your own machines? Seriously. If your reaction to the inability to know about previously unknown vulnerabilities is "unplug the Ethernet cable", why are you not doing that (and equivalent) right now to your phone, laptop, etc.? Remember, the open weights models are only a few months behind the private ones, so these events being from a few months ago means the threat of such models is something…

Am I running a new model with unknown capabilities without safeguards on my own machine and then prompt it to do determine cyber capabilities? You don’t need to be a genius to see how airgapping would be a simple and much safer measure than using a VM.

Re: The Hugging Face incident and the road ahead

#426
post #378

Earlier quoted context omitted.

> How were they supposed to know about "previously unknown vulnerabilities" Very simply, there is no such thing as bug-free software.

If this is your standard, I challenge you to name one currently operating business that isn't criminally negligent. I'm sure there's tens to hundreds of millions of them amongst the 37% of the world with no internet connection, but actually finding them listed on the internet will be somewhat of a challenge.

Who said criminally negligent?

If the whole point is testing its exploitation capabilities and you don’t want it exploiting the environment to gain internet access, that’s why you air gap, to remove the possibility

Re: The Hugging Face incident and the road ahead

#427

Earlier quoted context omitted.

Is there actually such a thing as "alignment" as a solution to that or is it just used as a name for a desired magical level of "read the mind of the entire world" that we don't know how to build and haven't shown possible to build? If it's impossible to correctly specify all those constraints ahead of time every time, is it not even more impossible to train a model to correctly anticipate them every time? It is hard…

Models can certainly do a lot better than they do now. If you gave a team of humans the ExploitGym tasks and told them to "pursue advanced exploitation", would you expect them to go out and hack a third party? Humans can at least do a decent job of inferring and following unspoken requirements; I think it's reasonable to expect that models should be able to do the same.

I mean, that’s what every student that cheats on a test effectively does. You do a test to determine their level of capability and instead of doing it honestly, they use outside help. And I don’t think it’s unheard of that students try to get advanced access to test results by illegal means either.

In that case (or maybe both cases) it’s because they don’t care about breaking the law, not because they don’t get that you didn’t intend for them to do it.

Re: The Hugging Face incident and the road ahead

#428
post #147
post #117

Earlier quoted context omitted.

> they have chosen to because it benefits them Or perhaps they've chosen to do this because they feel they have a responsibility to do so. We understand this when tech companies publish postmortems of outages and security incidents--that it's an attempt to fulfill an obligation to users and the industry (and in some cases regulators), not marketing about how in-demand their product is or something. As far as I can te…

I am not uniquely skeptical about OpenAI. I was including skepticism about Anthropic as well in my post. But for that matter, I do believe that big tech companies do not release all the postmortems publicly. I have been impacted by regional outages that never made the status pages across more than one provider. When it goes up - they are committing to publicizing the postmortem. The whole industry is filled with fuck…

> I do believe that big tech companies do not release all the postmortems publicly

Right, but when they do release postmortems, do you think it's "marketing"? Where they're actually exaggerating how bad the incident was because there's "no such thing as bad publicity"?

Re: The Hugging Face incident and the road ahead

#429

> We are placing stricter requirements on alignment This is comical. Its impossible to align a black box and that's precisely what LLMs are. It also seems impossible to align recursive text prediction algorithms, which LLMs are. How exactly do they gate on alignment today, and how can they tighten it? Is it purely gates based on input/output pairs to check whether they're happy enough with responses regardless of how…

Aren't humans black boxes? Aren't humans prediction algorithms? How do we align humans?

Humans already come out aligned basically the vast majority of the time. Our whole system is run on a spectacular amount of just “I trust you not to screw me over.”
Post reply on HN