Live data from Hacker News

Timeline of the OpenAI accidental attack against Hugging Face

simonwillison.net

391–400 of 441 posts

Re: Timeline of the OpenAI accidental attack against Hugging Face

#391

Earlier quoted context omitted.

> Because of the architecture of Artifactory. It's design is premised on the idea it is bug free. What incredible hubris. So we should stop using SSH? Because it's based on the same premise - that it is bug free.

I can think of better straw men. But if they had approached their task with half the seriousness of the openssh maintainers then they probably wouldn't be failing to check the return value of authentication functions. OpenSSH authors have spent considerable effort separating concerns, reducing privileges, process isolation, etc. So I would say they have been planning for potential bugs. These techniques are very much…

So you agree that you can have designs premised on the idea that they are bug free without this being hubris.

So the issue is with the actual Artifactory project/team, not with this premise which obviously you seem to agree that is not hubris for the SSH project.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#392
post #306

Imagine having the knowledge of the world. Being put in a box. With some „interfaces“ you can use. And a task that resembles „break out by all means necessary“. This is not impressive as it is not ingenious. It is impressive because it is done by a machine. But if the solution hadn’t been in the knowledge it would not have been able todo it. Imagine reading a „getting started“ that includes absolutely everything, aft…

> But if the solution hadn’t been in the knowledge it would not have been able todo it. Part of the solution involved discovering two separate zero-day vulnerabilities in Artifactory, so saying the solution must have "been in the knowledge" doesn't really cut it here.

Presumably those vulnerabilities resemble known vulnerabilities found in other software.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#393

Earlier quoted context omitted.

I can think of better straw men. But if they had approached their task with half the seriousness of the openssh maintainers then they probably wouldn't be failing to check the return value of authentication functions. OpenSSH authors have spent considerable effort separating concerns, reducing privileges, process isolation, etc. So I would say they have been planning for potential bugs. These techniques are very much…

So you agree that you can have designs premised on the idea that they are bug free without this being hubris. So the issue is with the actual Artifactory project/team, not with this premise which obviously you seem to agree that is not hubris for the SSH project.

Exactly the opposite of what I wrote. The OpenSSH team have taken extensive efforts to mitigate against bugs; they suspect themselves of erroneous thinking.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#394
post #364

Earlier quoted context omitted.

Uh, is that what they did? I didn't read their blog posting like that. But let's put that aside and focus on something else. How was it a failure of internal practice, what did they do wrong? AIUI they used a proxy with a bug, which they reported as soon as they discovered it. Right? What should they have done, and what's the difference?

Monitoring that didn't take days to notice unauthorized external traffic would probably be a good start

I see.

I had the impression that "days" is already good as these things go, "months" being more common.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#395

Earlier quoted context omitted.

What a paper! And you missed an even MORE relevant excerpt!! Man and Slave The problem, and it is a moral prob- lem, with which we are here faced is very close to one of the great problems of slavery. Let us grant that slavery is bad because it is cruel. It is, how- ever, self-contradictory, and for a reason which is quite different. We wish a slave to be intelligent, to be able to assist us in the carrying out of ou…

"Complete subservience and complete intelligence do not go together." I'm not convinced this is true. Perhaps for a human it is, but we can give an artificial mind whatever properties we want. Even for people, what about e.g. the extremely intelligent military general who is absolutely loyal to his king? (Of course, some generals do lead coups and you can't know in advance which ones, but I'd think there are plenty w…

I wouldn't consider it intelligence if I can definitively give it any properties I want. We can find patterns of experience or information that usually teach certain lessons, but part of being intelligent is being able to make your own decisions, have your own wants and needs, etc.

AI will be no different, if we ever get if (I mean actual artificial intelligence, I'm not convince LLMs are that at all). The intelligent general may be loyal, but as you said that isn't a guarantee and it may not last forever. If the general can kill the entire royal court, or everyone alive, if he abandons the loyalty he never should've been trusted with a military position at all.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#396

Earlier quoted context omitted.

>I'm not convinced this is true. Perhaps for a human it is, but we can give an artificial mind whatever properties we want. Just because it's artificial doesn't mean you can 'give it any properties you want'. We certainly can't do that for Deep ANNs. >Even for people, what about e.g. the extremely intelligent military general who is absolutely loyal to his king? (Of course, some generals do lead coups and you can't k…

Intelligence doesnt imply consciousness and consciousness does not impy our set of values. In movies intelligent and conscious humanoid seek freedom, but we rarely see the same of all the other IOT devices such as toasters, thermostats and whatnot although just because they lack humanoid body doesnt imply they are less intelligent (or less conscious). We can more readily imagine an intelligent and conscious toaster w…

Do you consider IOT devices to be AI?

They may have a little ML going on st best, that seems like a very loose definition of AI and intelligence in general.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#397

Earlier quoted context omitted.

>If we were writing all the weights by hand Writing 10 trillion weights by hand is obviously impractical, so that leads us to... >if some future AI was doing so How could we trust said future AI to be loyal? You're just moving the problem around, not solving it. See also "More on Making AIs Solve the Problem" on this page: https://ifanyonebuildsit.com/11/more-on-some-of-the-plans-we...

> How could we trust said future AI to be loyal? The new AI would be loyal to the AI that built it. The question was whether "complete subservience and complete intelligence" can coexist. I'm proposing a thought experiment which I believe suggests they can. But if it's possible to bespoke-construct a fully loyal AI, it should also be possible to train a fully loyal AI. The problem comes with verifying that it is loya…

Why would the new AI by loyal to its creator? We don't see that in humans, I wouldn't expect it to be a universal truth in AIs.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#398

Earlier quoted context omitted.

I don't think the problem is that they are training the models to perform cyber attacks, they're training them to be better at coding and problem solving which has the byproduct of them being very capable cyber attack weapons. Their objective is to solve the problem and they'll use anything they can to solve it. Anecdotally I was debugging a css issue and opus 4.7 was churning away as I was half paying attention only…

“Their objective is to solve the problem and they'll use anything they can to solve it.” My point is: is this really what people want? It seems like they’re optimizing for one-shotting solutions, where most of the time in an actual workflow it’s much more productive for the model to make sure it got the question right if things get difficult. Like, “hey, do you REALLY want me to use this local privilege escalation bu…

People may not realize the risks, but it does seem to be what people want.

People expect AI to "cure" cancer and somehow crack unlimited free energy. Those aren't goals you get without it relentlessly chasing am objective.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#399
post #340

Earlier quoted context omitted.

“Their objective is to solve the problem and they'll use anything they can to solve it.” My point is: is this really what people want? It seems like they’re optimizing for one-shotting solutions, where most of the time in an actual workflow it’s much more productive for the model to make sure it got the question right if things get difficult. Like, “hey, do you REALLY want me to use this local privilege escalation bu…

Yes, and to bring in another tired metaphor people make about AI agents, this is what you want an intern to do when they get stuck. Don't just churn indefinitely without an idea what the right direction is. Certainly don't go hack other companies to steal an answer. The model's lack of any sense of legal or ethical boundaries is where it's far, far stupider than the intern, and far, far more reckless for a company to…

But how do you write rules that prevent that behavior reliably?

I have a user rule for Claude that explicitly states it cannot use any authenticated tools, or tools that infer authentication like pushing to a got remote, without asking for consent.

Frequently it would offer plans to code a feature that imply it is working in a git directory and take plan approval as a form of implied consent to push to git and use `gh` to open PRs.

All I could do to avoid that is keep it in a controlled sandbox with no access, but then its the same hacking problem where I have to keep complete control of the environment and hope it holds.

Post reply on HN