Live data from Hacker News

Timeline of the OpenAI accidental attack against Hugging Face

simonwillison.net

251–260 of 302 posts

Re: Timeline of the OpenAI accidental attack against Hugging Face

#251
post #232
post #212

Earlier quoted context omitted.

I'm sure that's very easy to say when you financially benefit from it.

I expect I could make a whole lot of money blasting out sensationalist headlines about how the AI labs are all faking security incidents as part of their marketing campaigns.

I have to disagree. Someone who fully embraces and perpetuates sensationalist AI hype marketing like this would be far more likely to pay $10/mo to be fed more marketing than someone who questions and doubts it.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#253
post #251
post #232

Earlier quoted context omitted.

I expect I could make a whole lot of money blasting out sensationalist headlines about how the AI labs are all faking security incidents as part of their marketing campaigns.

I have to disagree. Someone who fully embraces and perpetuates sensationalist AI hype marketing like this would be far more likely to pay $10/mo to be fed more marketing than someone who questions and doubts it.

If someone wants to spend $10/month for exposure to sensationalist AI hype there are a whole lot of newsletters they should subscribe to that will deliver what they want better than I do.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#254

Earlier quoted context omitted.

What a paper! And you missed an even MORE relevant excerpt!! Man and Slave The problem, and it is a moral prob- lem, with which we are here faced is very close to one of the great problems of slavery. Let us grant that slavery is bad because it is cruel. It is, how- ever, self-contradictory, and for a reason which is quite different. We wish a slave to be intelligent, to be able to assist us in the carrying out of ou…

"Complete subservience and complete intelligence do not go together." I'm not convinced this is true. Perhaps for a human it is, but we can give an artificial mind whatever properties we want. Even for people, what about e.g. the extremely intelligent military general who is absolutely loyal to his king? (Of course, some generals do lead coups and you can't know in advance which ones, but I'd think there are plenty w…

You seem to be confusing intelligence with objective function.

Subservience seems to be sublimation of objectives to a master; intelligence seems to point out the ability to realize suboptimality of the master's objective function according to the master's actual objectives.

While an intelligent general may be absolutely loyal, he also would presumably help the king/president to avoid unproductive strategies.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#255
post #233

Earlier quoted context omitted.

> Usage of AI services has shifted dramatically to Chinese providers - from 4% at the beginning of the year to some 30% now. Where did you see that number?

I knew when I wrote that it was a bare assertion, based partly on memory. This is an approximation based on a few sources, the principal of which was this article, which pulls from a bunch of other sources in turn. https://www.secondtalent.com/resources/ai-trends-in-china/

Oh, it's the OpenRouter number: https://finance.yahoo.com/technology/ai/articles/china-ai-mo...

Those numbers aren't credible IMO because OpenRouter only see traffic for people who have chosen to route their traffic through OpenRouter. If you do that, you're much more likely to be experimenting with alternative models. They have no insight at all into people who point their applications directly at OpenAI or Anthropic without having OpenRouter in the middle.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#256
post #184

Earlier quoted context omitted.

Your comment is already showing the mistaken, poisonous belief of security maximalism, that tries to reinterpret_cast everything into hacks and cybersecurity vulnerabilities. Most of these things aren't "hacking". They're problem-solving and efficiently dealing with obstacles and random bullshit along the way. This , not "hacking", is what they're making their models "razor focused on". Problem is, most normal comput…

>They're problem-solving and efficiently dealing with obstacles They are problem solving as much as a falling rock is finding its path down a mountain.

I'd readily agree that they may be (probably are?) utterly unaware of what they're doing, with no spark of sapience.

However, I'm a sapient being employed as a software developer for my problem-solving ability.

If you gave me a Kobayashi Maru scenario as a challenge, I would probably come up with the idea of hacking out of the sandbox to find the answer.

If I was in a technical interview, I would probably even ask the interviewer if exploits are fair game, or if that's too far outside the box.

I highly doubt I'd find a new zero-day as quickly as these agents did.

I wouldn't say it's _impossible_ - I've found security issues before.

But I'm not a specialist, and I'd bet against myself.

If the agentic LLMs can consistently achieve something that's a bridge too far for me, then I don't know what to call that other than problem-solving.

I say this as an LLM hater who would push the "Nuke all LLMs" button the instant I had access to it.

Opus 4.8 and 5, at least, don't seem to me to be solving problems by deep, thorough understanding - my employers have compelled me to use Claude, so I've used them a lot to build things, and I constantly find both little and large hallucinations that scream "these are still missing something."

Maybe these new models are actually massively better, or maybe they're just the same kind of system 1 thinking done faster and harder.

The distinction is largely academic, though, for questions like "Can you keep these contained?", "Can you farm out arbitrary programming tasks to them and expect an acceptably mediocre answer?", or "Does it matter if these things are aligned?"

Re: Timeline of the OpenAI accidental attack against Hugging Face

#257

Earlier quoted context omitted.

"Car accidents occur therefore we shouldn't have cars" isn't very compelling.

It’d be more like “car accidents occur, so let’s add seat belts, air bags, etc…”.

... and speed limits

Re: Timeline of the OpenAI accidental attack against Hugging Face

#258

Earlier quoted context omitted.

What a paper! And you missed an even MORE relevant excerpt!! Man and Slave The problem, and it is a moral prob- lem, with which we are here faced is very close to one of the great problems of slavery. Let us grant that slavery is bad because it is cruel. It is, how- ever, self-contradictory, and for a reason which is quite different. We wish a slave to be intelligent, to be able to assist us in the carrying out of ou…

"Complete subservience and complete intelligence do not go together." I'm not convinced this is true. Perhaps for a human it is, but we can give an artificial mind whatever properties we want. Even for people, what about e.g. the extremely intelligent military general who is absolutely loyal to his king? (Of course, some generals do lead coups and you can't know in advance which ones, but I'd think there are plenty w…

The whole thing seems to depend upon AI agents objective ie to achieve some objective by any means possible and ignoring any guardrails. The article did not clarify if openAI had any guardrails to begin with while conducting this experiment. For all the talks around how much they invest in AI safety one would expect them to have these common sense guardrails in place or is it just a case of some school children letting their pet monkeys loose deliberately to display how awesome their monkey team is.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#259

I wish we could stop sensationalizing this about the AI and really just understand the incompetence of the labs disabling an internet connection in a sandbox.

As written it sounds like you're saying that it was incompetent of the labs to disable the sandbox internet access? They tried to disable open internet access but the models zero-day'd their Artifactory package registry and got internet access anyway. No sensation... that's just what happened.

Unplug the ethernet cable leading to the outside world, then?

Re: Timeline of the OpenAI accidental attack against Hugging Face

#260
post #99

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

Their position makes no sense to me. I don’t see how you can be a mainstream company selling your services worldwide (almost) if you also believe that you’re building an extremely dangerous AGI (supposedly based on the same technology you’re offering to everyone). If you actually believe that an AGI would be extremely dangerous that should 100% be a very strictly regulated area of research, similar to bio weapons. An…

I know some people who are worried at Anthropic, and their position seems to be "if we don't do it, someone even less responsible will. Unilateral disarmament didn't work and real oversight seems unlikely to happen in time, so we'll just try to be as safe as we can be (while still winning the race)"

Not that they're happy about it, they just see no other realistic choice

Post reply on HN