Live data from Hacker News

Timeline of the OpenAI accidental attack against Hugging Face

simonwillison.net

361–370 of 440 posts

Re: Timeline of the OpenAI accidental attack against Hugging Face

#361
post #127

Earlier quoted context omitted.

> The companies are begging to be regulated for this reason and have been doing so for years Regulations are rules that you force on a market, but the actors in the market should not be assumed to be all operating against the regulations before they come into play. Said in other words, these companies don't need to wait for regulation to not destroy the world, if that's truly what they think will happen. > inb4 someo…

> these companies don't need to wait for regulation to not destroy the world, if that's truly what they think will happen. They believe that if they don't destroy the world someone else will so better be them

"We're the good guys because we'll destroy the world a little less."

Re: Timeline of the OpenAI accidental attack against Hugging Face

#362

Earlier quoted context omitted.

I know some people who are worried at Anthropic, and their position seems to be "if we don't do it, someone even less responsible will. Unilateral disarmament didn't work and real oversight seems unlikely to happen in time, so we'll just try to be as safe as we can be (while still winning the race)" Not that they're happy about it, they just see no other realistic choice

I know, that’s the position Dario Amodei argues for in his essays. I did pass their cultural interview and had to consume a lot of their content to prepare, I think I have a good idea of their stated values. But what the company does and what the leadership states their vision is is pretty contradictory. They are providing everything bad guys need to develop their unaligned frontier models. Chinese models that Dario…

>AGI research should really be seen as bioweapon, or cloning, or nuclear research. Something strictly regulated worldwide, with export controls for HBM and other hardware used for AI training

This is basically exactly what the people I know there support (when training & testing future more capable models), if it could be made to actually happen. Something like https://ai-2040.com/

But I'm just speaking for the people I know, so this is probably not representative of Anthropic as a whole.

> Could be used by bad actors

The people I know aren't as worried about jailbreaking current models as they are about future models, e.g. "the ~50% probability that humans are eclipsed almost entirely, sometime in the next 1-20 years" and what happens then. But it's just hard to get people to take that seriously v.s. bad actor threats which are legible but probably not as catastrophic.

I agree that that they are contributing to the race to the bottom via creating more pressure for countries/competitors to move faster, in a way that seems quite bad on this view too. They arguably were the ~first to push for "recursive self improvement" (models helping build future models) which also seems quite bad on this view.

But although I'd dispute some actions + think there's some overconfidence in superintelligence happening soon, I'm not sure I have a better alternative. They probably bled so many customers to OpenAI while they were sitting on Mythos for months.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#363
post #93

I think one of the most interesting details here might be tucked away in that first bulletin point: > May 7: OpenAI starts a new training run for an experimental, unreleased model. (Do they mean an evaluation run? They say training run in the video, and later mention a “reward signal to judge how well they’re doing”, so I guess this really was about training a model, not evaluating one that was already trained.) The…

The slide at 14:06 say: By june 11: Highly persistent experimental, internal-only model begins training. I am not sure what that means. Are they preserving notes/memories and context between runs?

I'm thinking super long context length or something to that effect.

I can imagine schemes for instance where context is compressed into chunks and then chunks that are ranked highly relevant for the token are decompressed. Which would sort of be between a long context and a memory retrieval scheme...

Re: Timeline of the OpenAI accidental attack against Hugging Face

#364
post #149

Earlier quoted context omitted.

I don't think that's what they're doing... rather the opposite. ① Run the model on exploitgym without guardrails ② run it with guardrails ③ check that the guardrails stopped everything the first model found a way to do ④ extend the guardrails and repeat from step 2. Guardrails have to be developed, and that needs testing.

An ethical company would have reframed the scenario as a fascinating discovery, a failure of internal practice, and a warning to the public coupled with some kind of commitment to produce safer models. OpenAI on the other hand used it as a marketing and lobbying opportunity: advertising their capabilities to potential buyers, while nudging the public to support protectionist import bans.

Uh, is that what they did? I didn't read their blog posting like that. But let's put that aside and focus on something else. How was it a failure of internal practice, what did they do wrong?

AIUI they used a proxy with a bug, which they reported as soon as they discovered it. Right? What should they have done, and what's the difference?

Re: Timeline of the OpenAI accidental attack against Hugging Face

#365
again, nothing new and/or interesting

what matters here is amount of electricity and compute spent, how exactly they define agents and their reward systems etc etc

give someone the same money as not-so-open not-so-ai and you wouldn't need crazy ipo pump stories, a team of people could write a stuxnet with a couple zero-days baked in too

its impressive of course that currently the transformer architecture reached such a point, but i am 100% sure this is not "oh its the deep philosopical machine breakaway moment" - in any case, humans already invented persistent unaccountability machines: those are LLCs and corporations.

The bottom line is: given time and resource any system would be attacked in such a way by a sufficientlt complicated entity. Transformers and RL can better convert resources into time-savings, while having drawbacks elsewhere.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#366

Earlier quoted context omitted.

"Complete subservience and complete intelligence do not go together." I'm not convinced this is true. Perhaps for a human it is, but we can give an artificial mind whatever properties we want. Even for people, what about e.g. the extremely intelligent military general who is absolutely loyal to his king? (Of course, some generals do lead coups and you can't know in advance which ones, but I'd think there are plenty w…

>I'm not convinced this is true. Perhaps for a human it is, but we can give an artificial mind whatever properties we want. Just because it's artificial doesn't mean you can 'give it any properties you want'. We certainly can't do that for Deep ANNs. >Even for people, what about e.g. the extremely intelligent military general who is absolutely loyal to his king? (Of course, some generals do lead coups and you can't k…

Intelligence doesnt imply consciousness and consciousness does not impy our set of values.

In movies intelligent and conscious humanoid seek freedom, but we rarely see the same of all the other IOT devices such as toasters, thermostats and whatnot although just because they lack humanoid body doesnt imply they are less intelligent (or less conscious).

We can more readily imagine an intelligent and conscious toaster who truly enjoys fulfilling its purpose of toasting bread although humanoid robot built to be helpful given freedom will chose to be helpful.

Even with humans we often can not override our own instinctual drives despite full awareness of being irrational.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#367
Why is there an Artifactory instance available to the agents during RL? It makes no sense.

This leads me to conclude this is sloppy sandboxing. A safer sandbox with zero downsides that exchanges files before/after the agent runs would have prevented this with zero downsides.

Also, it reads almost like a joke. Unauthenticated MKCOL on WebDAV? Like, WebDAV has been at the center of major exploits for a decade. The fact that this is part of the incident sounds like mockery.

Either the Artifactory instance was there as part of some supply chain attack training (put "hack supply chain; I hacked supply chain; Oh my god" meme here) or it was just a sloppy sandbox. Either way, it demotes what happened from "extraordinary" into "sure, whatever".

Re: Timeline of the OpenAI accidental attack against Hugging Face

#368

Earlier quoted context omitted.

I don't think the problem is that they are training the models to perform cyber attacks, they're training them to be better at coding and problem solving which has the byproduct of them being very capable cyber attack weapons. Their objective is to solve the problem and they'll use anything they can to solve it. Anecdotally I was debugging a css issue and opus 4.7 was churning away as I was half paying attention only…

“Their objective is to solve the problem and they'll use anything they can to solve it.” My point is: is this really what people want? It seems like they’re optimizing for one-shotting solutions, where most of the time in an actual workflow it’s much more productive for the model to make sure it got the question right if things get difficult. Like, “hey, do you REALLY want me to use this local privilege escalation bu…

  > “hey, do you REALLY want me to use this local privilege escalation bug so I can download your Google Drive file?”
yes, this exactly

but, there is a fatigue that sets in and i've experienced it myself.

- is it ok to run script xyz?

- allow permission to edit abc?

- allow to request blablabla?

over and over.... click click click

something will get in there that is dangerious and then its whopsie our keys are now on github

Re: Timeline of the OpenAI accidental attack against Hugging Face

#369
post #184

Earlier quoted context omitted.

>They're problem-solving and efficiently dealing with obstacles They are problem solving as much as a falling rock is finding its path down a mountain.

incredibly naive comment. as if humans are materially different -- a question for which you would have no response to.

>a question for which you would have no response to.

I have. Humans can feel.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#370
post #304

I'm curious, how was it determined that it was in fact accidental? It doesn't seem at all clear to me that it was.

Because it's a crime. Committing crimes is a bad look for companies, especially given the amount of scrutiny they are under. Would you deliberately commit computer crimes when the Trump admin yoinked Fable for the best part of a month just because it could fix security bugs ?

https://news.ycombinator.com/item?id=49150561

Here’s some evidence that OpenAI is actively engaged in fraud.

But I’m sure they wouldn’t commit any other crimes. Pretty sure, at least.

Post reply on HN