Live data from Hacker News

Timeline of the OpenAI accidental attack against Hugging Face

simonwillison.net

301–310 of 441 posts

Re: Timeline of the OpenAI accidental attack against Hugging Face

#301

Earlier quoted context omitted.

"Complete subservience and complete intelligence do not go together." I'm not convinced this is true. Perhaps for a human it is, but we can give an artificial mind whatever properties we want. Even for people, what about e.g. the extremely intelligent military general who is absolutely loyal to his king? (Of course, some generals do lead coups and you can't know in advance which ones, but I'd think there are plenty w…

>I'm not convinced this is true. Perhaps for a human it is, but we can give an artificial mind whatever properties we want. Just because it's artificial doesn't mean you can 'give it any properties you want'. We certainly can't do that for Deep ANNs. >Even for people, what about e.g. the extremely intelligent military general who is absolutely loyal to his king? (Of course, some generals do lead coups and you can't k…

> We certainly can't do that for Deep ANNs

Only because we don't know how! We don't actually understand how weights work, so we make computers come up with the weights instead. If we were writing all the weights by hand--or if some future AI was doing so--why couldn't we make it perfectly loyal?

Re: Timeline of the OpenAI accidental attack against Hugging Face

#302
post #37

Earlier quoted context omitted.

> But the event itself only seems possible because they failed to properly monitor and isolate the environment in the first place. OpenAI is clearly run by dummies and subpar engineering talent. > The model is obviously impressive Speak for yourself.

Speaking of that "obviously impressive" line, I'm getting really tired of something like that line seemingly needing to be included by anyone doing any criticism of agentic systems. The most common form of it is "these models are obviously useful" midway through a bunch of arguments about environment, data provenance, skill atrophy, or even correctness issues. It's just really weird. Why does everyone feel the need t…

This may be out of left field but you might be interested in Michael Parenti's essay "left-wing anti communism". It's about this same thing in American politics where everyone from the furthest right to furthest left has to condemn socialism before opening their mouth, and how it's turned the US's elected left into preemptively apologetic losers.

No idea where you stand politically but there's not that many arguments about this type of rhetorical error so hopefully you consider it.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#303
post #46

Norbert Wiener in 1960: "As is now generally admitted, over a limited range of operation, machines act far more rapidly than human beings and are far more precise in performing the details of their operations. This being the case, even when machines do not in any way transcend man's intelligence, they very well may, and often do, transcend man in the performance of tasks. An intelligent understanding of their mode of…

"Car accidents occur therefore we shouldn't have cars" isn't very compelling.

A while ago I noticed that car crashes were the leading cause of death for age ranges too old for infant mortality and too young for heart failure.

I checked again before making this reply and found that in many cases "accidental poisoning" has overtaken car crashes. Accidental poisoning is overwhelmingly "drugs".

I do find your argument compelling even if you do not.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#304

I'm curious, how was it determined that it was in fact accidental? It doesn't seem at all clear to me that it was.

Because it's a crime. Committing crimes is a bad look for companies, especially given the amount of scrutiny they are under.

Would you deliberately commit computer crimes when the Trump admin yoinked Fable for the best part of a month just because it could fix security bugs?

Re: Timeline of the OpenAI accidental attack against Hugging Face

#305
Imagine having the knowledge of the world. Being put in a box. With some „interfaces“ you can use. And a task that resembles „break out by all means necessary“.

This is not impressive as it is not ingenious. It is impressive because it is done by a machine. But if the solution hadn’t been in the knowledge it would not have been able todo it.

Imagine reading a „getting started“ that includes absolutely everything, after that all is just like a set of Lego, given enough time you will have what is asked for. But nothing original, because it never had an original thought.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#306

Imagine having the knowledge of the world. Being put in a box. With some „interfaces“ you can use. And a task that resembles „break out by all means necessary“. This is not impressive as it is not ingenious. It is impressive because it is done by a machine. But if the solution hadn’t been in the knowledge it would not have been able todo it. Imagine reading a „getting started“ that includes absolutely everything, aft…

> But if the solution hadn’t been in the knowledge it would not have been able todo it.

Part of the solution involved discovering two separate zero-day vulnerabilities in Artifactory, so saying the solution must have "been in the knowledge" doesn't really cut it here.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#307

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

I don't think the problem is that they are training the models to perform cyber attacks, they're training them to be better at coding and problem solving which has the byproduct of them being very capable cyber attack weapons. Their objective is to solve the problem and they'll use anything they can to solve it. Anecdotally I was debugging a css issue and opus 4.7 was churning away as I was half paying attention only…

A tool that will "do anything they can to solve it" including illegal and unhelpful things does not seem like a good tool to me.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#308

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

Your comment is already showing the mistaken, poisonous belief of security maximalism, that tries to reinterpret_cast everything into hacks and cybersecurity vulnerabilities. Most of these things aren't "hacking". They're problem-solving and efficiently dealing with obstacles and random bullshit along the way. This , not "hacking", is what they're making their models "razor focused on". Problem is, most normal comput…

What is your definition of hacking if it doesn't include using leaked security tokens scraped from the web? Also, kirk 100% cheated.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#309

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

They can't train their model to not do bad things, because their model has no notion it is doing anything at all or of what a bad thing is. It's only predicting the next token, and in doing so producing a facsimile of intelligence. The best they can do is create guardrails, which will only work probabilistically. In other words, those guardrails will fail at certain points on the probability curve. Of course that's n…

Yeah so this falls into the engineering trap of "well it's hard so we can skip that part."

If they can't train things safely then they shouldn't do it at all.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#310

From the outside, it looks like OpenAI got exactly the kind of event they could market the hell out of to demonstrate the capability of the model. But the event itself only seems possible because they failed to properly monitor and isolate the environment in the first place. To me, it looks like their job is to market the model, not take security seriously. The model is obviously impressive, but we already knew that.…

I'm quite sure the whole event is planned. Not planned in a sense that OpenAI employees carefully designed every step, but in a sense that ignoring security practices was desired and intentional. >> Show me the incentive and I'll show you the outcome. Once you realize security breaches are marketable, a security breach is just around the corner.

I agree - also kinda funny that Meta followed and also reported a breach by their model, "They are getting PR, lets do the same!"
Post reply on HN