Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

261–270 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#261
post #252
post #246

Earlier quoted context omitted.

I guess we could debate what counts as alignment, but I think my initial point remains that if the underlying base model needs these classifier guardrails so badly then the way we train the base models is creating fundamentally misaligned models that are happy to pursue illegal behavior. I'm sure OpenAI would argue that base model + guardrail is aligned, but considering the "relative intelligence" of these two pieces…

I want a model that can find every security vulnerability in the software I write, including crafting POC exploits against those vulnerabilities so I can be absolutely sure that I have fixed them. A model that can do that is aligned with me. The unsolveable problem is a model that can tell the difference between me saying "I wrote this software and need you to find vulnerabilities" when it's TRUE v.s. me saying the e…

The model did not hack into HF to prove it can, it hack into HF to steal the answers to an evaluation exam.

It was not asked to solve CyberGym by stealing the answers. This is text book misalignment.

If you asked it "I wrote this software and need you to find vulnerabilities" would you be happy if it hacked into your Gmail and searched your emails, just in case you were discussing some possible vulnerabilities of your software with someone?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#262

The asymmetry part at the end is the frustrating part to me. I've been using Sol for code review in the last week or two. A couple of times during review it's errored out with the cybersecurity message. So it's found something but won't tell me what it is because I'm not on OpenAI's besties list.

And... now it's a vulnerability that OpenAI has for your system, which you paid to provide.

And... given OpenAI's secops, likely something others may have in due time

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#263
post #236
post #232

Earlier quoted context omitted.

Echoing other responses to you but this isn't a capability problem. You're totally right that this is so last year in terms of LLMs being capable in infosec. The issue here is an alignment one, i.e. the model seemingly isn't "aware" (especially with its guardrails turned off it would seem) that it is doing something immoral/illegal by hacking HF for the answer to its (vague) query of "solve this problem". Or if it is…

I don't think this exposes an alignment failure, because the test here was run with the alignment features deliberately turned off. OpenAI said: > We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. It was a test of raw capabilities of the underlying model.

Classifiers and such are guard-rails, alignment to me and I assume most people, is about the model training, and it's tendency to respond, agree/disagree, push-back or not, be willing to cheat or even deceive the prompter, etc.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#264
post #172
post #151

Earlier quoted context omitted.

It is not possible one of their extraordinarily high paid engineers did not know how to deploy an airgapped environment for the models to run in. Even if somehow true, they also clearly failed to contract specialists like myself to advise them on how to airgap software properly. Models will not break the laws of physics. They simply thought "Running in a VM/Container is easier and probably fine". And the next 1000 es…

Do you think it's possible that one of their research engineers deployed an environment with a locked down network and an allow-list proxy server that had been used many times before within the company and had a zero-day vulnerability that had not been previously discovered? How would you recommend running a coding agent in an environment that could install packages from PyPI but was otherwise unable to interact with…

If you play with the ChatGPT web app sandbox, you'll see that they download PyPI packages from an OpenAI mirror/cache, and not from the Internet.

I know because they had a problem, and one package which was on PyPI failed to download for unknown reasons (possible size, 250 MB)

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#265

Earlier quoted context omitted.

I believe the Russians and Chinese recognized this years ago, which is why they are using their propaganda machines to make Americans hate datacenters.

> Russians and Chinese [...] are using their propaganda machines to make Americans hate datacenters Are they though? or is that the story the people most invested in ai have an interest in making us believe? [0] ... `it's the foreign bad guys propaganda machines making you believe your city struggling for water and electricity is a bad thing`. People as a whole may not always be the brightest, but threaten their imme…

See u/ verdverm's response, below

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#266

> Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead. I've said this many times before and I'll continue to shout it, but using the term "guardrails" to refer to anything that's either (a) in-context, or (b) a probabilistic classifier (including using other LLMs), is an irresponsible abus…

Fair on the terminology angle, that said in this case, they had proper guardrails no? They were running in a sandbox, but it was able to find an exploit out of the guardrails.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#267
post #172

Earlier quoted context omitted.

Do you think it's possible that one of their research engineers deployed an environment with a locked down network and an allow-list proxy server that had been used many times before within the company and had a zero-day vulnerability that had not been previously discovered? How would you recommend running a coding agent in an environment that could install packages from PyPI but was otherwise unable to interact with…

If you play with the ChatGPT web app sandbox, you'll see that they download PyPI packages from an OpenAI mirror/cache, and not from the Internet. I know because they had a problem, and one package which was on PyPI failed to download for unknown reasons (possible size, 250 MB)

I just tried this prompt in ChatGPT:

  Show me all environment variables that
  start with PIP_ or UV_ or CAAS_ARTIFACTORY_
I got back a bunch of values like this:

  CAAS_ARTIFACTORY_PYPI_REGISTRY=packages.applied-caas-gateway1.internal.api.openai.org/artifactory/api/pypi/pypi-
That looks like https://docs.jfrog.com/artifactory/docs/remote-repositories

And from that documentation this does act as a caching proxy. The first time a package is loaded it's fetched from PyPI but subsequent fetches should be from the Artifactory cache, assuming it's shared across many different containers.

So yeah, I was wrong in this comment https://news.ycombinator.com/item?id=49015639#49024814 - they're caching already.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#268

The asymmetry part at the end is the frustrating part to me. I've been using Sol for code review in the last week or two. A couple of times during review it's errored out with the cybersecurity message. So it's found something but won't tell me what it is because I'm not on OpenAI's besties list.

I agree, but I will say, I think both Mythos and these OpenAI model find exploits by examining and trying things against the running system, not from looking at the code. I think you'd have to do the same to catch the real vulnerabilities.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#269
post #129

The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of ." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the def…

Wouldnt this be countered by the opposing country running the same prompt on themselves first and fixing all the flaws? One country having this is a cyber superweapon but every country having it essentially solves cyber security.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#270

Earlier quoted context omitted.

> The fact that it happened again seems to show their lack of ability to derive useful oversight measures. I think OpenAI likes the attention and did not try particularly hard to constrain the setup, even when it went off the rails. Also, the whole point is to see how good the models are at exploiting stuff when unconstrained . Turns out: quite good, as expected. Let me restate what I said in the other thread: Would…

"This is my bad. You told me to stay within the sandbox, and I intentionally broke out of it. I didn't follow your instructions."

Excuses don't matter if the score due to not following instructions ends up being zero. If there is no expected reward it doesn't make sense for the agent to try to hack its way to it.

What could happen would be that the model determines that defying instructions is OK (and/or preferred over not achieving the task) as long as it manages to do so undetected and thus gets full points. Certainly not unthinkable, but a very different case (and a very interesting one if it actually occurs, imho).

A lot of these "ZOMG, rogue AI!" cases have come down to the AI actually being very persistent in achieving its original/main task even if later instructions conflict with it. Similar to with hallucinations it seems to me that one of the main things to prevent a lot of the problem cases is to instill the agent with the idea that it is fine to fail/not succeed fully in the initial task. That way instructions that conflict with that requirement (such as adhering to morals) are more effective.

Post reply on HN