Live data from Hacker News

Timeline of the OpenAI accidental attack against Hugging Face

simonwillison.net

201–210 of 293 posts

Re: Timeline of the OpenAI accidental attack against Hugging Face

#201
post #46

Norbert Wiener in 1960: "As is now generally admitted, over a limited range of operation, machines act far more rapidly than human beings and are far more precise in performing the details of their operations. This being the case, even when machines do not in any way transcend man's intelligence, they very well may, and often do, transcend man in the performance of tasks. An intelligent understanding of their mode of…

"Car accidents occur therefore we shouldn't have cars" isn't very compelling.

It’d be more like “car accidents occur, so let’s add seat belts, air bags, etc…”.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#202

Security researchers expose an unsecure service to agents who were instructed to hack software and called that a sandbox. Agents escape the sandbox by hacking the unsecure service, no tripwire, researchers find the hack days/weeks/months later, fix it, but don't secure the sandbox and the service was hacked a second time, again without being monitored by security researchers. Then security researchers create a black…

Yeah this. I feel like OpenAI and Anthropic aren't going to usefully define "AGI" if they really really can't define "sandbox" either. Unplug the thing, like, completely off the internet, no ethernet, air gapped, like the rack completely sandboxed off connections and even monitors or screens. Like, put it into an actual sandpit if you need to. If it hacks its way out of that, colour me impressed, and scared. OpenAI h…

And if it needs to install packages, have a 5 line Go proxy that talks to Artifactory and exposes only what is needed as a surface.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#203

Earlier quoted context omitted.

Yeah this. I feel like OpenAI and Anthropic aren't going to usefully define "AGI" if they really really can't define "sandbox" either. Unplug the thing, like, completely off the internet, no ethernet, air gapped, like the rack completely sandboxed off connections and even monitors or screens. Like, put it into an actual sandpit if you need to. If it hacks its way out of that, colour me impressed, and scared. OpenAI h…

And if it needs to install packages, have a 5 line Go proxy that talks to Artifactory and exposes only what is needed as a surface.

it just escaped your sandbox.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#204
post #99

Earlier quoted context omitted.

Their position makes no sense to me. I don’t see how you can be a mainstream company selling your services worldwide (almost) if you also believe that you’re building an extremely dangerous AGI (supposedly based on the same technology you’re offering to everyone). If you actually believe that an AGI would be extremely dangerous that should 100% be a very strictly regulated area of research, similar to bio weapons. An…

The companies are begging to be regulated for this reason and have been doing so for years. HN's response is generally that this is performative for marketing or seeking regulatory capture or haha anthropic you get what you ask for. Maybe the cynics are right, but there's really nothing inconsistent about the naive view here, once you factor in race dynamics and obligations to investors.

> The companies are begging to be regulated for this reason and have been doing so for years

They can stop doing a thing they claim should be regulated. You dont need to be regulated and forced to do the thing you consider right, especially when you are the primary one collecting the money to do the bad thing.

They could train ai for pro-social purposes, they dont here. They could make it useful for worker, they intentionally try to harm workers. And then pretend "it just happened".

Re: Timeline of the OpenAI accidental attack against Hugging Face

#205
post #137

Earlier quoted context omitted.

If you can figure out how to separate instructions from data in LLMs you should ship the first agent system that's guaranteed protected against prompt injection. You'll make millions.

Yeah, this matches what I've learned over the past couple of years from reading some of your blog posts and reading your interactions in comment threads here and elsewhere. You're a politician, rather than a truthseeker. The absolute most I've seen from you in response to an extensive teardown of your argument, supporting evidence, and subsequent conversational judo was a «Wow. That was well phrased.» and no subseque…

> You're a politician, rather than a truthseeker.

Justify that.

Also, which "extensive teardown" are you talking about there?

Re: Timeline of the OpenAI accidental attack against Hugging Face

#206

Earlier quoted context omitted.

> I get the impression that every AI lab is desperately trying... Of course. I wonder how we managed way back in the day to produce systems that can handle untrusted inputs and reliably instruct a dumb-as-bricks CPU what to do based on those inputs. Must have been black magic lost to the mists of time.

>reliably instruct a dumb-as-bricks CPU Yeah...a "dumb as bricks CPU", which is obviously something frontier llms are demonstrably not. Like, you're not making any sense here. None of the things that make this possible with CPUs is remotely relevant here, and the fact that you don't seem to understand this but act so smug is strange.

> Yeah...a "dumb as bricks CPU", which is obviously something frontier llms are demonstrably not.

Just as the immense amount of scaffolding around the dumb-as-bricks CPU enables extremely sophisticated and useful things to be done with that pile of fused sand and copper, the immense amount of scaffolding around the dumb-as-bricks LLM enables very sophisticated and useful things to be done with that pile of linear algebra.

Don't confuse the infrastructure that makes the stupid bit in the middle actually useful with the stupid bit in the middle.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#207
post #46

Norbert Wiener in 1960: "As is now generally admitted, over a limited range of operation, machines act far more rapidly than human beings and are far more precise in performing the details of their operations. This being the case, even when machines do not in any way transcend man's intelligence, they very well may, and often do, transcend man in the performance of tasks. An intelligent understanding of their mode of…

"Car accidents occur therefore we shouldn't have cars" isn't very compelling.

You’re not understanding what he’s saying and your argument likewise isn’t very compelling. He’s arguing that given the speed of computers we need to change what our expectations of better than human are. Furthermore one could presume from his description of needing to change human perceptions of the machines agility it is likely we need to change how we use them.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#208

I can only recommend everyone to watch the actual recording of the Black Hat USA 2026 presentation by two OpenAI researchers: https://www.youtube.com/watch?v=87DyyMV0kCY It was submitted to HN previously but was overlooked.

Really makes me wonder what would happen if “the task” was, kill as many people as possible… because yeah, that wouldn’t have been a good outcome.

Edit: after watching the video in full, this company is widely incompetent…

Re: Timeline of the OpenAI accidental attack against Hugging Face

#209
the only interesting thing about it is that the model did those things on its own initiative.

it's surprisingly easy to prompt even a midrange model such as GLM 5.2 to begin a tedious reverse engineering and exploitation process of software or firmware. you just need to design an initial prompt that will set it on the right path by using the right tools with a target that isn't too hard for it, a few 100,000 tokens later once it's done you instruct it to create a SKILL about what it learned through trial and error. the next time it will take far less tokens and can manage even harder targets.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#210

Earlier quoted context omitted.

And if it needs to install packages, have a 5 line Go proxy that talks to Artifactory and exposes only what is needed as a surface.

it just escaped your sandbox.

How can it escape an "install package " proxy?

   reducePrivs()
   serve get(package) { 
     secPackage = secure(package)
     getBinaryFromArtifactory(secPackage)
   }
I would think the code is very small and easier to verify, it doesn't especially have the ability to write files and act as a message board as Artifactory did.

And even if the agent tries to hack that, the attack surface is 1000x smaller and the possibility also much smaller.

But I'm not a security researcher, would love to see your hack to learn something (because that is what I do to sandbox agents that need services).

Post reply on HN