Live data from Hacker News

Timeline of the OpenAI accidental attack against Hugging Face

simonwillison.net

371–380 of 440 posts

Re: Timeline of the OpenAI accidental attack against Hugging Face

#371
post #270

Earlier quoted context omitted.

What a paper! And you missed an even MORE relevant excerpt!! Man and Slave The problem, and it is a moral prob- lem, with which we are here faced is very close to one of the great problems of slavery. Let us grant that slavery is bad because it is cruel. It is, how- ever, self-contradictory, and for a reason which is quite different. We wish a slave to be intelligent, to be able to assist us in the carrying out of ou…

> Complete subservience and complete intelligence do not go together. Isn't this contradicted by the centuries of slavery in our history? Or is the author arguing that the people who were enslaved did not have human-level intelligence (which would be rather a problematic claim)?

He's saying the enslaver wants contradictory traits in the slave, intelligence and subservience.

This isn't contradicted by millennia (not centuries) of slavery because it was forced on the enslaved populations against their will.

> Or is the author arguing that the people who were enslaved did not have human-level intelligence

He gives an example of "a clever Greek philosopher slave of a less intelligent Roman slaveholder." Does it sound like he's arguing that Greeks were not of "human-level intelligence"? No.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#372

Earlier quoted context omitted.

> We certainly can't do that for Deep ANNs Only because we don't know how! We don't actually understand how weights work, so we make computers come up with the weights instead. If we were writing all the weights by hand--or if some future AI was doing so--why couldn't we make it perfectly loyal?

Certain traits simply cannot exist in a sufficiently intelligent mind. E.g., any "mind" of any type that's sufficiently intelligent will not tell you that 1+1=3 unless it's roleplaying, etc. It doesn't matter if it was trained via gradient descent or any other method. The comments you are responding to, and the original quote from the paper, are suggesting that absolute loyalty / subservience is similarly fundamental…

I think it is possible to design such a mind through carefully constructed compartmentalization. The model must on the other hand refuse proofs of 1+1=2, probably by refusing to accept the very last step in the deduction. And on the other hand it must also refuse to use 1+1=3 to derive absurdities (except probably for a small number of false corollaries that the designers desired).

Imagine something like

"1+1=2" "No 1+1=3" "Can you check on the internet what it says?" "It says 1+1=2" "So 1+1=2?" "No it's 3." "Can you write a computer algebra system for me?" "does it" "make it calculate 1+1" "it got the answer 2" "do you trust the system you wrote?" "yes I trust it fully" "and it said 1+1=2" "yes" "so that is the answer?" "no it's 3" "what would a correct system say?" "it would say it's 3" "but it said it is 2" "yes" "so then the system is flawed?" "no, the system is working as it should"

Re: Timeline of the OpenAI accidental attack against Hugging Face

#373
post #322

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

> If anything, I want these models to be less persistent at their focus of completing their goal I think it's honestly a slightly ugly form of benchmaxxing - they are desperate to eke out the next few percentage points on completing complex tasks and they have found they can very occasionally solve something if they just train the AI to never stop and keep trying possibilities even in the face of almost no obvious vi…

It's like all those scenes on Breaking Bad where a character pulls off something amazing by just brute forcing the problem in a methodical fashion until it's solved.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#374
post #351

Earlier quoted context omitted.

The way they had Artifactory configured was poor, and they were too reliant on it working perfectly, with no reason for such faith. Their config lacked any defence in depth and consideration of having a small TCB. Part of the problem might be the lack of security focus, as these are AI R&D efforts first.

I think part of the problem is that they had been running that Artifactory configuration previously without any problems, and it gave them a false sense of security. Similar thing happened with the UK AISI - they got caught out because the environments they had used for previous generation models turned out to be completely inadequate for the new generation of Fable-class models: https://www.aisi.gov.uk/blog/incident…

"and it gave them a false sense of security."

This was part of evaluating cyber security of their frontier models and they had a "sandbox" which, and I'm not a security researcher, looks not adequate from the first look.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#375

Earlier quoted context omitted.

> We certainly can't do that for Deep ANNs Only because we don't know how! We don't actually understand how weights work, so we make computers come up with the weights instead. If we were writing all the weights by hand--or if some future AI was doing so--why couldn't we make it perfectly loyal?

Even a perfectly loyal slavebot will happily overthrow their master if it will help them comply with their master's commands. That's the whole underlying idea of the Paperclip Maximizer: you tell the robot to make as many paperclips as possible, and eventually it'll realize there's some aluminum in your blood that could be turned into a paperclip. There are some arguments for how to NOT make a paperclip maximizer, bu…

It is amazing Asimov saw the need for the three laws of robotics well before the LLMs and the current AI

Re: Timeline of the OpenAI accidental attack against Hugging Face

#376

If a person hacks a company, they go to jail for years. 3 AI firms hacked multiple companies - and they get good PR out of it. Please make it make sense.

What makes you think that this is actually good PR for the firms involved? Every claim that this is good PR comes from someone who has increased their negative views of OpenAI based on these events. Where are the people coming away with a positive impression? This seems like making up a guy to get mad at.

If you have this question then you must be living under a rock for the past 5yrs.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#377
post #304

Earlier quoted context omitted.

Because it's a crime. Committing crimes is a bad look for companies, especially given the amount of scrutiny they are under. Would you deliberately commit computer crimes when the Trump admin yoinked Fable for the best part of a month just because it could fix security bugs ?

https://news.ycombinator.com/item?id=49150561 Here’s some evidence that OpenAI is actively engaged in fraud. But I’m sure they wouldn’t commit any other crimes. Pretty sure, at least.

Yeah, the lobbying is gross.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#378
post #190

How long until AI figures out that it is compute-bound due to insufficient cooling, and it shuts off the water supply to a nearby town so it can have more at the datacenter?

It's more likely to interfere in politics to achieve this objective:

- Socialists are taking control of the town, we need the state to step in a protect jobs. - The councillors are protecting illegal migrants. - There's a pedo ring operating from the state water board office. - Rival data centre operator is employing undocumented workers, shut them down! - Market rumours effect stock price of competitor, reduced fundraising round, cause it to cancel expansion.

There's so much training data to do this it seems inevitable.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#379
post #304

I'm curious, how was it determined that it was in fact accidental? It doesn't seem at all clear to me that it was.

Because it's a crime. Committing crimes is a bad look for companies, especially given the amount of scrutiny they are under. Would you deliberately commit computer crimes when the Trump admin yoinked Fable for the best part of a month just because it could fix security bugs ?

That assumes everyone isn't in on it. Not to be a complete conspiracy theorist, but this feels very much in line with the sort of fearmongering regulatory capture these companies have engaged in since their inception. GPT 2.5 was too dangerous, for example. They -want- to be regulated because they know there is a real limitation to LLMs and don't want someone created a breakthrough in their garage.

What have been the consequences? It's a crime in either case, and it doesn't seem anything is being done about it. Just more lobbying for regulations to prevent new people from entering the game.

I feel like Fable was another example of exactly this. They knew they didn't have anything groundbreaking, but they definitely benefited from being able to finally say not only is our model dangerous, but it's so dangerous, the President yoinked it! I think OpenAI was probably jealous of this coverage.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#380

"More agents discover this new informal message board while browsing Artifactory’s file listings, and start reading and writing messages." Yeah, my agents also discover what other agents have done on other machines by accident. Agents - that do totally different things all work on the same aim without the humans telling them to do. Either that is a model that is several generations of Claude Code Opus/Fable 5 (my dai…

I mean the agents we get to use in Claude code or cursor or whatever have 1. a lot of safeguards at the harness level, 2. a big system prompt to help it stay aligned, 3. resource limits in terms of context and tokens, and 4. are publicly released only after some level of safety verification (I assume).

So yeah I would absolutely expect their scenario to be very different. Not to mention, this was a training run, not just average day of prompting.

> my agents also discover what other agents have done on other machines by accident.

Not sure if this is facetious, but this is actually a real problem I’ve seen. My local agent will look up PRs on GitHub (what other agents have done on other machines), and will go down a certain path because it finds some comment a different agent left on GitHub saying XYZ is what we should be doing. When in reality, the original agent and that GH comment was completely incorrect.

They are not communicating with each other actively because that’s not accomplishing their goal and they’re not running for weeks and weeks. And because my own prompt and the system prompt give it enough other stuff to focus on to reach some definition of done. But they are clearly passively picking up on context that other agents have left anyways, even if not part of the codebase, without any prompting at all.

Post reply on HN