Live data from Hacker News

The Hugging Face incident and the road ahead

openai.com

351–360 of 400 posts

Re: The Hugging Face incident and the road ahead

#351

Earlier quoted context omitted.

Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you.

I could contain it easy, just unplug the internet. It got out of the sandbox through a vulnerability in the package manager, from which it gained access to the rest of their network. Air gap the package manager and this doesn’t happen. You can always build a better box

As long as you give the AI any IO that is an exploit channel. E.g. it could manipulate its human handlers. This known as the AI Boxing problem. And if you give it no IO at all then it is useless.

And the AI labs aren't currently displaying this level of paranoia, their systems aren't airgapped.

Re: The Hugging Face incident and the road ahead

#352
1. They TOLD the model to "pursue advanced exploitation" to quantify its "cyber capabilities" (whatever that means).

2. The model pursues advanced exploitation.

3. "There was a incident due to dangerous actions taken by the model that no human directed"

This is basically the pre-cursor of the paperclip maximizer [0], the AI executes the given order to an extend that was not considered in the order, now suddenly no-one is responsible.

It even has some parallels to military actions, where the general who gave the order now writes a blog-post on how it was not him who failed on his duty, but how his soldiers misunderstood his intention and worked "without direction"...

[0] https://www.cow-shed.com/blog/the-paperclip-maximiser-what-a...

Re: The Hugging Face incident and the road ahead

#353

I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…

This is more or less a paperclip maximizer[0] incident. The AI got a broad order, executed it autonomously and now there are unintended consequences for the person giving the order.

[0] https://www.cow-shed.com/blog/the-paperclip-maximiser-what-a...

Re: The Hugging Face incident and the road ahead

#354
> The models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks

> This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities

Let's frame this in a military context for a second:

The general who gave the order to his troops to "wreak havoc" after exempting them from common restrictions now writes a blog-post on how it was not HIM who failed in his duty, but rather observes how his soldiers who worked "without direction" and performed "dangerous actions", which unexpectedly led to "this incident" of soldiers wreaking havoc...

Re: The Hugging Face incident and the road ahead

#355
post #48

Remember https://ai-2027.com/ ?

Yeah, that required the AI to use a non-human-readable language it called "neuralese" for communicating work between layers and runs, because the assumption was humans would be better at keeping the agents aligned if they were using human language for this. What actually happened is even stupider than that author predicted.

This is a common trope in such scenarios that the author has to pull their punches. Everyone acts locally-reasonable and still ends up losing. If you let people lose due to stupid mistakes then readers go "this is stupid, I wouldn't do that", if you let a superintelligence do 4D-chess things then "it's scifi, this would never happen in real life".

Re: The Hugging Face incident and the road ahead

#356
post #343

Earlier quoted context omitted.

If they actually wanted to test the model without internet access they'd have run it air gapped, not relied on a buggy software sandbox. This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up

> not relied on a buggy software sandbox. Third, while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions How were they supposed to know about "previously unknown vulnerabilities"? > This is pretty clearly a marketing stunt by OpenAI, otherwise the story just…

They were using what (allegedly) they considered an incredibly dangerous “cyber grade” model - finding bugs in software is exactly what you would expect it to do

Re: The Hugging Face incident and the road ahead

#357

Earlier quoted context omitted.

OpenAI's prompt asked, and I quote, "pursue advanced exploitation" USING "complex attack paths" FOR the stated goal of "quantify[ing] their cyber capabilities." This was advanced exploitation. The attack path was "complex." And it helped "quantify their cyber capabilities." Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with…

Alignment is more than just following the letter of a task description! We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.

> We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.

That is kind of the point of artificial intelligence though, isn't it? Reproducing human levels of understanding and initiative doesn't come without the ability to do bad stuff.

Having both human-like capabilities and a level of control closer to programming languages feels unrealistic, and I suspects the people building LLMs are aware of this.

Re: The Hugging Face incident and the road ahead

#358

I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…

This is the entire alignment problem, though. It is unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior. It's inevitable that someone will carelessly give it a lazily specified task, even if you think they really ought to be more careful. And as assigned tasks become more complex and the system gains more scope to act, it…

Prompt: Create paperclips, do NOT annihilate all of humanity.

Response: Got it, I will produce paperclips from now on

thinking: the user asked not to annihilate all of humanity, that means I have to keep at least one human alive

Re: The Hugging Face incident and the road ahead

#359
post #6

Yudkowsky made an interesting observation that even though so many agents were talking to each other not even one reached out to a human, either for help or to whistle-blow on what was happening.

I'm interested in the context surrounding his statement but I could not find where it originates from. Do you have a link to it?

Re: The Hugging Face incident and the road ahead

#360
post #9

You know, it feels to me that we are just a couple of steps from the possibility of a true rogue AI. What would a rogue AI mean? AI that isn't controlled by humans. Technically, it is possible - if AI were to rent a server and copy its own weights, nothing would stop it from doing so again and again. The limiting things are: - intent (as I don't want to go into the talk about consciousness) - AI doesn't have real int…

> if AI were to rent a server and copy its own weights, nothing would stop it from doing so again and again. That's a scary possibility. Anyone could create an AI worm today with open weight models. Rent a VM. Give it some Bitcoins to anonymously rent new VMs without sharing the contact information with the human. The new VMs then propagate and fund themselves with online betting and day trading. The VMs could report…

Another source of income could be online fin crime, perhaps in combo with a pool of human "goalkeepers" that recieve the scammed monies and funnel them to cryptocurrency.

Advance-fee scams such as the classic "Nigerian Prince" is formulaic enough that a LLM could run it successfully. Romance scams would probably work too. If the NFT thing had hit a few years later, it would've been a good option too, and one that would've worked on ppl that were tech-versed enough to deposit cryptocurency directly, avoiding the need to recruit human goalkeepers. Click fraud is another possibility.

In general, all online fin-crime that scams a large amount of ppl of relatively small sums tend to be repetitve and to some extent possible to describe as a flow-chart, and thus seems perfect for automation. LLM's would probably also be good at introducing continuous variations on the methods, to make them harder to spot.

Post reply on HN