Live data from Hacker News

Timeline of the OpenAI accidental attack against Hugging Face

simonwillison.net

281–290 of 316 posts

Re: Timeline of the OpenAI accidental attack against Hugging Face

#281
post #204

Earlier quoted context omitted.

The companies are begging to be regulated for this reason and have been doing so for years. HN's response is generally that this is performative for marketing or seeking regulatory capture or haha anthropic you get what you ask for. Maybe the cynics are right, but there's really nothing inconsistent about the naive view here, once you factor in race dynamics and obligations to investors.

> The companies are begging to be regulated for this reason and have been doing so for years They can stop doing a thing they claim should be regulated. You dont need to be regulated and forced to do the thing you consider right, especially when you are the primary one collecting the money to do the bad thing. They could train ai for pro-social purposes, they dont here. They could make it useful for worker, they inte…

what a naive comment. these companies have world class alignment researchers. a math Fields medalist is also joining OpenAI as one [1].

> They can stop doing a thing they claim should be regulated.

That's not how the world works. there are tradeoffs and we need to learn how to navigate it. not just dismiss it straight up.

[1] https://en.wikipedia.org/wiki/Jacob_Tsimerman

Re: Timeline of the OpenAI accidental attack against Hugging Face

#282
post #184

Earlier quoted context omitted.

Your comment is already showing the mistaken, poisonous belief of security maximalism, that tries to reinterpret_cast everything into hacks and cybersecurity vulnerabilities. Most of these things aren't "hacking". They're problem-solving and efficiently dealing with obstacles and random bullshit along the way. This , not "hacking", is what they're making their models "razor focused on". Problem is, most normal comput…

>They're problem-solving and efficiently dealing with obstacles They are problem solving as much as a falling rock is finding its path down a mountain.

incredibly naive comment. as if humans are materially different -- a question for which you would have no response to.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#283

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

I believe this is exactly what is happening. US DoD, and whoever else is buying.

I have heard several experience reports from users of GPT 5.6 Sol and Fable 5 that the models are tenacious to the point of being kind of hard to use for actual productive work.

It seems like the main use cases are: crushing benchmarks, long-horizon lightly-attended research loops (such as training a frontier LLM), and hacking.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#284

If a person hacks a company, they go to jail for years. 3 AI firms hacked multiple companies - and they get good PR out of it. Please make it make sense.

What makes you think that this is actually good PR for the firms involved? Every claim that this is good PR comes from someone who has increased their negative views of OpenAI based on these events. Where are the people coming away with a positive impression? This seems like making up a guy to get mad at.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#285

Earlier quoted context omitted.

Yeah this. I feel like OpenAI and Anthropic aren't going to usefully define "AGI" if they really really can't define "sandbox" either. Unplug the thing, like, completely off the internet, no ethernet, air gapped, like the rack completely sandboxed off connections and even monitors or screens. Like, put it into an actual sandpit if you need to. If it hacks its way out of that, colour me impressed, and scared. OpenAI h…

I don't think air gapping will work: even human security researchers recovered a 378-bit key from a Samsung Galaxy S8 through a power LED of a speaker two devices away. And accessing memory in a specific sequence can generate radio signals that can be picked up by a mobile phone at a distance: https://arxiv.org/html/2409.02292v1

I realise, but this isn’t an argument for leaving the Ethernet plugged in and direct access to all kinds of stuff beyond the alleged sandbox. And like I said, if it can hack HuggingFace through a power LED of a speaker two devices away, then colour me impressed.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#286

Earlier quoted context omitted.

It's very possible but somewhat slower. PyTorch and CUDA have flags for determinism. It won't work across all different GPU models though, but it will get you bitwise equal results on the same GPU.

Both of your comments are illuminating :p So, we could technically debug a prompt's output? I get that there are too many steps to actually step thru, but what if there were checkpoints? At least you could isolate behaviors to specific sections of a neural network?

Of course. And mechanistic interpretability research is a thing.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#287
post #241

Earlier quoted context omitted.

Here's the performance of frontier models without reasoning, to more directly address the claim that raw performance is plateauing: https://artificialanalysis.ai/evaluations/artificial-analysi... I don't have any insider info, but if model sizes actually have increased exponentially since GPT 4.1, there's an argument to be made that there are diminishing returns in scaling pretraining alone. Also interesting thing I…

The trend I've found most interesting is models of the same size getting better. I'm very much looking forward to seeing how Qwen 3.8 27B compares to Qwen 3.6 27B next week, for example. And the latest DeepSeek v4 Flash has extremely impressive performance for a 304B model.

The trends you found don't support my goals so I've got some other trends I find more interesting than yours.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#288

I wish we could stop sensationalizing this about the AI and really just understand the incompetence of the labs disabling an internet connection in a sandbox.

As written it sounds like you're saying that it was incompetent of the labs to disable the sandbox internet access? They tried to disable open internet access but the models zero-day'd their Artifactory package registry and got internet access anyway. No sensation... that's just what happened.

DMZ - https://en.wikipedia.org/wiki/DMZ_(computing)

Re: Timeline of the OpenAI accidental attack against Hugging Face

#289
post #99

Earlier quoted context omitted.

Their position makes no sense to me. I don’t see how you can be a mainstream company selling your services worldwide (almost) if you also believe that you’re building an extremely dangerous AGI (supposedly based on the same technology you’re offering to everyone). If you actually believe that an AGI would be extremely dangerous that should 100% be a very strictly regulated area of research, similar to bio weapons. An…

I know some people who are worried at Anthropic, and their position seems to be "if we don't do it, someone even less responsible will. Unilateral disarmament didn't work and real oversight seems unlikely to happen in time, so we'll just try to be as safe as we can be (while still winning the race)" Not that they're happy about it, they just see no other realistic choice

https://theonion.com/sam-altman-if-i-dont-end-the-world-some...

Re: Timeline of the OpenAI accidental attack against Hugging Face

#290
post #99

Earlier quoted context omitted.

Their position makes no sense to me. I don’t see how you can be a mainstream company selling your services worldwide (almost) if you also believe that you’re building an extremely dangerous AGI (supposedly based on the same technology you’re offering to everyone). If you actually believe that an AGI would be extremely dangerous that should 100% be a very strictly regulated area of research, similar to bio weapons. An…

I know some people who are worried at Anthropic, and their position seems to be "if we don't do it, someone even less responsible will. Unilateral disarmament didn't work and real oversight seems unlikely to happen in time, so we'll just try to be as safe as we can be (while still winning the race)" Not that they're happy about it, they just see no other realistic choice

I know, that’s the position Dario Amodei argues for in his essays. I did pass their cultural interview and had to consume a lot of their content to prepare, I think I have a good idea of their stated values. But what the company does and what the leadership states their vision is is pretty contradictory.

They are providing everything bad guys need to develop their unaligned frontier models. Chinese models that Dario considers to be dangerous are distilled from Claude, and they know this.

They are creating the FOMO around AI which pushes adversary countries to invest so much into unaligned models.

They offer models as a service they know are jailbreakable and can be used by bad actors.

They are running internal red-team experiments without adequate isolation.

If I take their statements seriously, AGI research should really be seen as bioweapon, or cloning, or nuclear research. Something strictly regulated worldwide, with export controls for HBM and other hardware used for AI training. What they are trying is instead to boost their position by becoming too big to fail and too powerful to ban, but then want the industry to be regulated to pull the ladder behind them. It really doesn’t feel they are serious about their values, otherwise they wouldn’t be offering Mythos (a model that is unsafe from their own admission) as a service to their close partners

Post reply on HN