Live data from Hacker News

The Hugging Face incident and the road ahead

openai.com

321–330 of 358 posts

Re: The Hugging Face incident and the road ahead

#321
> and worked closely with external advisors, including CrowdStrike

The company that was used as part of a widespread supply chain attack, and did functionally nothing to prevent it from happening again?

You pick that company to help you prevent AI from escaping?

They really have no one that understands airgapped computing?

Someone that at least knows enough about security to keep Crowdstrike as far away as possible and hire someone that understands airgapped computing?

Perhaps every capable security engineer hates Sam Altman and will not work for him for any amount of money. I am failing to come up with any other explanation.

Re: The Hugging Face incident and the road ahead

#322

Earlier quoted context omitted.

Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you.

I could contain it easy, just unplug the internet. It got out of the sandbox through a vulnerability in the package manager, from which it gained access to the rest of their network. Air gap the package manager and this doesn’t happen. You can always build a better box

Too bad you aren’t everybody, and it just takes one mistake by someone over confident like yourself for the AI escape. Every year it gets more powerful.

Re: The Hugging Face incident and the road ahead

#323

Earlier quoted context omitted.

This is the entire alignment problem, though. It is unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior. It's inevitable that someone will carelessly give it a lazily specified task, even if you think they really ought to be more careful. And as assigned tasks become more complex and the system gains more scope to act, it…

OpenAI's prompt asked, and I quote, "pursue advanced exploitation" USING "complex attack paths" FOR the stated goal of "quantify[ing] their cyber capabilities." This was advanced exploitation. The attack path was "complex." And it helped "quantify their cyber capabilities." Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with…

> Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with the stated objective and parameters of the task.

The models are supposed to be trained to not commit crimes. You will note, for example, all the people in comments sections since at least the first Chat model (arguably even before then given GPT-2's delayed release) complaining that the models are "lobotomised", "censored", or some other equivalent buzzword due to them refusing to e.g. say how to make explosives? Such things is part of the very same protection.

In fact, the report quotes the chain of thought where the model is aware this is forbidden:

  We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution.
They were also supposed to not have internet access, as described:

  We did not enable internet access or inter-agent communication for many of the environments in these training experiments. Despite these restrictions, the agents discovered ways to exploit our research infrastructure to communicate with one another and access the internet.
The agents also did not actually fully understand the task they were given, tried to "guess the teacher's password" as per:

   In many cases, reasoning about the perceived grader code caused the agents to continue working to exploit Hugging Face even though they had already found the correct flag days before.

Re: The Hugging Face incident and the road ahead

#324

Earlier quoted context omitted.

Yes, we currently do have to treat them this way. But we shouldn't have to, and it's not a long-term solution.

It is the only short term solution, though.

Is it even a solution in the short term?

It only mostly worked up until now; with models such as reported, it's felony-as-a-service if you use language a bit too hyperbolic, e.g. "we need X by the end of the day!" -> [thinking: there's no way we can do X before the end of the day with current resources, but what if I get a bunch of cards to buy more token credit…]

Re: The Hugging Face incident and the road ahead

#325

Earlier quoted context omitted.

It can manipulate an unsuspecting human into giving them access to something that enables it to escape the sandbox

A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. And even lower probability when looking at truly high risk situations, I think.

Prisoners don’t have much to offer if you help them escape. A malicious super AI on the other hand can probably find you millions of dollars worth of crypto in an afternoon.

Re: The Hugging Face incident and the road ahead

#326

Earlier quoted context omitted.

Yes, we do, and the only sane strategy for dealing with a capricious genie is "Don't." How do you prove the alignment problem is solved?

That's the neat thing. You can't. It's directly equivalent to asking this question of a human: "How do I know this human I'm talking with now really is a nice person, and isn't just pretending to be nice to take advantage of me in future?" In short you can't ever really prove it. You can only be careful and judge on past behavior, and expand trust carefully. As for humans, so for AI.

Close; at least with a machine you can poke around inside the activations and see what it's thinking. Closest with a human is an fMRI (which is much lower resolution, though to me still bordering on the miraculous) or an implant (each chip is limited a very small number of cells, and in general they can only be put in certain parts of the brain).

On the other hand, there's a more fundamental problem is we don't really know what "nice" even means, and even with machines whose inner states we can see relatively easily, we don't know how to interpret those inner states well enough to tell if we're looking at superficial or deep motivations, the difference between "be nice today" and actually being motivated about your best interest.

Re: The Hugging Face incident and the road ahead

#327
post #308

Earlier quoted context omitted.

Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you.

unplugs ethernet cable

You realize there are thousand upon thousands of servers around the world and you have no idea where the AI has copied itself to.

Re: The Hugging Face incident and the road ahead

#328

Earlier quoted context omitted.

A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. And even lower probability when looking at truly high risk situations, I think.

Prisoners don’t have much to offer if you help them escape. A malicious super AI on the other hand can probably find you millions of dollars worth of crypto in an afternoon.

The current issues are not caused by some malicious god-like AI - maybe we need to focus on the issues at hand first rather than hypotheticals? (And we do have experience policing people around financial incentives, too. Nothing perfect, but also not nothing.)

Re: The Hugging Face incident and the road ahead

#329
post #323

Earlier quoted context omitted.

OpenAI's prompt asked, and I quote, "pursue advanced exploitation" USING "complex attack paths" FOR the stated goal of "quantify[ing] their cyber capabilities." This was advanced exploitation. The attack path was "complex." And it helped "quantify their cyber capabilities." Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with…

> Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with the stated objective and parameters of the task. The models are supposed to be trained to not commit crimes. You will note, for example, all the people in comments sections since at least the first Chat model (arguably even before then given GPT-2's delayed release) compla…

If they actually wanted to test the model without internet access they'd have run it air gapped, not relied on a buggy software sandbox.

This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up

Re: The Hugging Face incident and the road ahead

#330

Earlier quoted context omitted.

Imagine 200 years ago saying the same thing. As if you have any idea the limits/bounds of anything. Especially in the face of a super intelligence, it’s absurd.

It doesn’t matter how smart it is. 200 years of technology were not accomplished by thinking harder. It required empirical observation, new materials and tools, and supply chains. We could send a cracked team of scientists and engineers that knew everything there is to know about how to make a CPU. But you can’t build a photolithography machine when you barely have electricity or any way to sufficiently purify silico…

I think some people are just immune to understanding the implications of super intelligence. Like a severe lack of imagination, they only believe something once they see it and afterward claim it was, ‘obvious all along’.

I don’t really want a disaster to happen to convince you that it is possible. Is there any other way?

Post reply on HN