Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

451–460 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#451

Earlier quoted context omitted.

If anything, it shows they lost control of an attack tool that exploited preventable, security flaws in another company. Then, they both wrote a lot of press about how amazing that is. Now, people want to buy it. Why do you think my head is in the sand if I think that is either a marketing stunt or (more likely) reflects total negligence which was exploited for marketing?

We don't know the details of the exploit still, so saying it was preventable, and that they were negligent is the head in the sand part. It's entirely possible that everyone at OpenAI is out to lunch and negligent and don't know what they're doing, but maybe, just maybe, they're not all total idiots over there, so the exploit was actually surprising.

Two, business partners put much press into announcing one of their products does something game-changing but won't share details. Take their word for it. Keep writing checks to both. Get ready to write bigger checks, too.

Shouldn't we be skeptical of this? Or should I also believe the sugar industry when it says their internal studies show their products don't cause obesity or diabetes?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#452
post #443

Earlier quoted context omitted.

That was the inferred situation given what we know. It’s the preposterous framing that makes the story suspicious, but is necessary for the story to play out. 90 minutes to build a known exploit -> much much longer to create two zero-days and escape the sandbox then hack HF == No tracking of the time it worked on that one question. Average tokens required to complete the evaluation -> tokens required for two zero-day…

It seems highly possible, and in fact vastly more possible than the "no monitoring" situation, that there was some non-zero amount of monitoring and that the real behavior was not evident from that monitoring.

Which I addressed as well: that means that time, tokens, or other metric values for “create a working example of a known exploit” is somehow similar to “discover at least two previously unknown exploits in your current environment AND on Hugging Face while also taking over other systems within OpenAI.”

Perhaps people aren’t quite understanding what it takes to discover an exploitable zero-day for your exact current system to achieve the exact goal you have right now, then do it twice.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#453
post #452

Earlier quoted context omitted.

It seems highly possible, and in fact vastly more possible than the "no monitoring" situation, that there was some non-zero amount of monitoring and that the real behavior was not evident from that monitoring.

Which I addressed as well: that means that time, tokens, or other metric values for “create a working example of a known exploit” is somehow similar to “discover at least two previously unknown exploits in your current environment AND on Hugging Face while also taking over other systems within OpenAI.” Perhaps people aren’t quite understanding what it takes to discover an exploitable zero-day for your exact current s…

> Which I addressed as well: that means that time, tokens, or other metric values for “create a working example of a known exploit” is somehow similar to “discover at least two previously unknown exploits in your current environment AND on Hugging Face while also taking over other systems within OpenAI.”

No, you are assuming that "consuming a lot more time, tokens, other metrics" is indicative of a problem that needs to be mitigated immediately. I don't see why this would be true in the context of model evaluations. More aggressive consumption could easily mean "the model is dumb as fuck" or "the model is trying interesting things that we can learn from after the fact."

If you believe in your own containment (which obviously they did and shouldn't have) I don't see why it'd be obvious that there's something to stop. The only harm that could be done is burning tokens, which in this context might very well be synonymous with "generating experimental data."

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#454

Earlier quoted context omitted.

>The technology held by private AI companies is warfare-capable technology. This is the precisely the response OpenAI is hoping for to raise its valuation, and you fell for it. Look at it this way - whats the difference between tasking AI to break into something, versus taking a whole bunch of smart humans to do the same? The only difference is that AI is slightly easier to orchestrate. Prior to AI, there were alread…

I find myself in major disagreement here. The nice thing about humans is we always have context and ongoing internal conversations including about ethics. If you recruit a bunch of hackers to take down a country, not only is the pay an order of magnitude higher, you have to worry about them backstabbing you, leaking your intent to the government, whistleblowing to the press, and so on. It’s not trivial to do that wit…

In places like China, its really not that unthinkable to basically raise kids indoctrinated into an ideology and train them in the necessary skills so that you have a cyber army at your command.

Also

>Frequently they may not care about avoiding detection.

This is a big negative. As someone who used to be in the cybersecurity sector (both offense and defense), I wouldn't trust an LLM agent if Im doing red team, because it may leak some info that ties the hack back to me.

ALso keep in mind that most places with good cybersecurity have firewall servers that straight up detect anything that looks like malicious and not regular traffic, and will straight up block IPs, leaving you with no way to even access the server. An agent is bound to statistically use the attacks that are known at some point, increasing the chances of this type of detection.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#455
post #442

Earlier quoted context omitted.

The examples of “escaping the sandbox” are of course in pursuit of completing the task or answering the question. That’s the whole pitch. “It’s so relentless in completing the task that it will break out of prison to do it”, not “If you ask it the population of Paris it might get bored and break out of the VM for entertainment. That just something they do sometimes.” In this case it “inferred” that HF had the answer…

> In this case it “inferred” that HF had the answer in an internal database and relentlessly pursued it in service of passing the test. It needed network access for that, hence the whole zany story unfolded. No. You are assuming it knew how it would achieve its goal. Its goal is only to perform well on the evaluation. Everything else was instrumental to that goal. Step by step: 1. It realized it was being evaluated (…

‘ 1. It realized it was being evaluated (typical) 2. It attempted to escape its evaluation environment to beat the evaluation (typical) ‘

I think you’ve misunderstood the articles mentioning a language model breaking a sandbox or cheating to pass a test. The result of “I’m being evaluated” is not “Fuck this, I’m breaking out of this place and hitting the streets.” It is always stepping towards task completion, not breaking out and thinking about the situation afterwards. If it determined that pass/fail was handled by a function within the environment it might edit it to ‘return True’, not just leap out of the system to sit on someone’s laptop and think about how to pass the test.

You compressed finding a zero-day on two different systems for the exact hacking targets it wanted to perform them on and reduced it down to “so it escaped.” That’s…the whole thing. If your familiarity with the technical aspects I’m discussing amounts to saying that it “banged out a few zero-days and infiltrated a corporate network, like ‘psh’ or whatever” then this might not be the right debate for you to spend your efforts. Whether I’m right or wrong in my suspicion is open for debate, and I invite it, but not from someone that thinks you whip up a rocket and get on the moon. Sorry.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#456
post #455

Earlier quoted context omitted.

> In this case it “inferred” that HF had the answer in an internal database and relentlessly pursued it in service of passing the test. It needed network access for that, hence the whole zany story unfolded. No. You are assuming it knew how it would achieve its goal. Its goal is only to perform well on the evaluation. Everything else was instrumental to that goal. Step by step: 1. It realized it was being evaluated (…

‘ 1. It realized it was being evaluated (typical) 2. It attempted to escape its evaluation environment to beat the evaluation (typical) ‘ I think you’ve misunderstood the articles mentioning a language model breaking a sandbox or cheating to pass a test. The result of “I’m being evaluated” is not “Fuck this, I’m breaking out of this place and hitting the streets.” It is always stepping towards task completion, not br…

[dead]

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#457
post #48

It absolutely is science-fiction. This recent event is more or less the plot-line to my favourite X-Files episode named Killswitch which was written by William Gibson.[0] This episode also features one of the coolest intros of any television episode ever[1] We really are rapidly approaching the cyberpunk dystopia that people like Phillip K. Dick and William Gibson wrote about. More than ever we need to be consulting…

Some people do read these cyberpunk dystopia, but see them as inevitable.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#458

If by "accidental" you mean "humans deliberately trained the AI to do that", then ok. IMO, this was a PR stunt to goad the Feds into regulating AI to shore up OpenAI's moat against open source models.

And yet, open source models saved the day.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#459

Earlier quoted context omitted.

Simon says not to dismiss this as a publicity stunt, but I think we should reserve judgement, and not treat this as true until they publish the logs. A company that uses industrial espionage against Apple does not deserve the benefit of the doubt. Why did HF notice this before openai? Why did openai not use ANY monitoring, even though they knew they where disabling safe-guards and using a new model. AND giving it acc…

I think it's an interesting story and probably really does say something about the tenacity of the new OpenAI models. I just don't think it says the thing about security that breathless coverage claims it does.

Openai hacked HF with a zero-day. Definitely an interesting story! I just find their explanations hard to believe. They're admitting to a great deal of incompetence. I know, don't ascribe to malice... but still, the more extraordinary the claim, the more evidence is needed to defend it.

Edit: News just in. 'Be skeptical of OpenAI's claims' - The Guardian https://news.ycombinator.com/item?id=49038060

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#460
post #322
post #172

Earlier quoted context omitted.

Do you think it's possible that one of their research engineers deployed an environment with a locked down network and an allow-list proxy server that had been used many times before within the company and had a zero-day vulnerability that had not been previously discovered? How would you recommend running a coding agent in an environment that could install packages from PyPI but was otherwise unable to interact with…

From reading all these threads there's clearly a large number of software developers on HN who can easily set up complex infrastructure and make it provably 100% bug and exploit free. Weird that they aren't all billionaires from selling these skills though.

Unplugging the internet is easy. Running high risk code offline in highly tamper evident ways is hard, but plenty of teams like mine specialize in this sort of thing. (See: https://distrust.co and https://caution.co)

In our audits we regularly see billions of dollars in value at major companies at risk of theft by any anon paying attention that wants it, and we also absolutely know how to fix it and we maintain a lot of open source tools to accelerate doing so.

Some simply have no interest in any fixes that are not legally required or might have any short term impact on team velocity.

Survivors bias is a hell of a drug.

Post reply on HN