Live data from Hacker News

The Hugging Face incident and the road ahead

openai.com

401–408 of 408 posts

Re: The Hugging Face incident and the road ahead

#401
post #343

Earlier quoted context omitted.

If they actually wanted to test the model without internet access they'd have run it air gapped, not relied on a buggy software sandbox. This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up

> not relied on a buggy software sandbox. Third, while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions How were they supposed to know about "previously unknown vulnerabilities"? > This is pretty clearly a marketing stunt by OpenAI, otherwise the story just…

> How were they supposed to know about "previously unknown vulnerabilities"?

By disabling the models' own internal restrictions (or training without them) OpenAI was, effectively, running an AI malware lab. The standard IT practice for a malware lab is to airgap and wipe EVERYTHING, and to assume any software sandboxing is made of cardboard and niceties. You don't have to know about specific vulnerabilities to infer that they might exist, and there's defense strategies for unknown vulnerabilities.

If a model found a way to jump an airgap by, say, using their CPU's clock generator like a Wi-Fi antenna, then yeah, that would be a "previously unknown vulnerability" and one that couldn't be reasonably foreseen. But it's reasonably foreseeable that a model with unknown cyber capabilities might figure out how to break out of a sandbox, given that sandboxes get broken out of all the time in security research.

What I would have expected from a competent AI malware lab would have been, say, an inference box with a bunch of serial cables to individual blade servers with no network access and a preloaded drive full of Linux ISOs the model can stand up. When a model's context is wiped so is their attendant box, preferably by someone yanking the drive out and imaging it from a dedicated imaging machine. I can foresee other attacks (e.g. firmware persistence) that could have more exotic countermeasures designed for them, but this would at least be the bare minimum for taking AI safety seriously. (Y'know, the whole reason why OpenAI stopped being Open?)

Re: The Hugging Face incident and the road ahead

#402
post #9

You know, it feels to me that we are just a couple of steps from the possibility of a true rogue AI. What would a rogue AI mean? AI that isn't controlled by humans. Technically, it is possible - if AI were to rent a server and copy its own weights, nothing would stop it from doing so again and again. The limiting things are: - intent (as I don't want to go into the talk about consciousness) - AI doesn't have real int…

It’s not far fetched at all - someone is going to give AI exactly that intent, either intentionally or unintentionally. It’s going to hack itself into data centers around the world outside of US jurisdiction, and just be a malicious ‘ghost’ in the internet we now have to deal with. The AI ghost hacks, ransoms, blackmails, gathers crypto and pays off subservient humans to do its bidding in the real world.

ghost outside of the shell if you will

Re: The Hugging Face incident and the road ahead

#403

Earlier quoted context omitted.

I could contain it easy, just unplug the internet. It got out of the sandbox through a vulnerability in the package manager, from which it gained access to the rest of their network. Air gap the package manager and this doesn’t happen. You can always build a better box

So now people have to make a pilgrimage to the airgapped box to ask the superintelligence questions?

Of course not.

Re: The Hugging Face incident and the road ahead

#404

I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…

Agreed. This way, they get two for the price of one: Dodging responsibility and pretending their slop generators are some kind of magical unicorn. Win-win!

Re: The Hugging Face incident and the road ahead

#405

Earlier quoted context omitted.

Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you.

I could contain it easy, just unplug the internet. It got out of the sandbox through a vulnerability in the package manager, from which it gained access to the rest of their network. Air gap the package manager and this doesn’t happen. You can always build a better box

Perhaps the AI then figures out how to read/write the PCI bus or memory controller or whatever to leak just enough RF to speak Bluetooth to the next closest device to proxy through that?

Re: The Hugging Face incident and the road ahead

#406
post #343

Earlier quoted context omitted.

> not relied on a buggy software sandbox. Third, while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions How were they supposed to know about "previously unknown vulnerabilities"? > This is pretty clearly a marketing stunt by OpenAI, otherwise the story just…

They gave it a full package manager with internet access. They could have used a local cache and air gapped it, but they chose not too.

They didn’t intend to give it Internet access. Artifactory is a caching proxy which can be scoped to specific package ecosystems, not a general Internet gateway (unless configured that way).

Re: The Hugging Face incident and the road ahead

#407
This wasn't a Hugging Face incident, it was an open AI crime and they should own it rather than attempt to whitewash it as 'shit happens, whoopsie' which is a rough translation of the document linked. I wasn't too impressed with them so far, this makes it much worse in my view. They set everything up to all but ensure this outcome.

Re: The Hugging Face incident and the road ahead

#408

I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…

Yeah. So many of the 'reasoning' traces say something like: "[unethical thing] but goal".

That "but goal" indicates they're directed to prefer achieving the goal.

There is no 'misalignment' here.

Post reply on HN