Live data from Hacker News

Timeline of the OpenAI accidental attack against Hugging Face

simonwillison.net

131–140 of 287 posts

Re: Timeline of the OpenAI accidental attack against Hugging Face

#131
post #105

Why are the agents trying so hard to communicate with each other, leaving messages and so on?

We are assigning semantics to systems that deal only in syntactics. The entire problem with the current "AI" hype is squarely based on how we interpret output from systems based on statistical modelling of natural language.

That software is built on top of human language and these systems can be used for uncanny automation is a huge societal problem at the moment because we are all assigning meaning to patterns that inherently have none. It's all just bits flicking back and forth. We can make them match human language and use such systems to store and process data for us. We can use these bits to turn equipment on and off and run physical systems in factories and so laboratories. And now we can use GPU farms to dazzle us with output streams that might look a lot like autonomous agents capable of understanding human language and automating computer tasks.

The failure modes, the so-called "hallucinations", the amount of model whispering going on in managing "harnesses", "instructions" and so on... It's all just a lot of confusion and pareidolia.

We should never have hooked up hospitals and water supply systems to the internet but now here we are: people can type text such as "find vulnerabilities and get access blah blah" into a box and it goes into a looping interaction with statistical models of language and out come streams of commands that some python parses and runs like a script kiddie into some virtual machine running kali linux and that may disrupt vital infrastructure...

None of that was inevitable, or necessary. None of that means anything. There is no genie in the GPU farm. We concocted this entire shadow theater and are collectively gasping as the marionette slices the throat of some guy in the front row. Who had the brilliant idea of tying the sharpened sword to the marionette and sit people within range?

Why did we plug everything into the academic network built on trust? Why did we build GPU farms and interactive loops getting them to produce commands that we then parse and run blindly in internet connected vms?

The entire thing has cost hundreds of billions of dollars so far and counting. And why? Because the mountains of shitty saas code has become too boring to work on? We have made software so garish that we cannot bear to work on it without these contraptions helping us fling code at wall at industrial levels? Substitute corporate-speak and -bureaucracy for software to extend to the rest of the economy.

This entire state of things is comical.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#132
post #107

Simon's retelling is more compact but it also invites anthropomorphization of the sharing of the familiarity with the message board which re-emerged a few times. Zvi's retelling handles this better. Zvi speculates that the secret message board familiarity was carried because it had been trained into the May-and-subsequent models: https://thezvi.substack.com/p/openai-trained-its-models-for-...

Seems like an artifact of the subagent pattern which is explicitly included in recent models.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#134
post #127

Earlier quoted context omitted.

The companies are begging to be regulated for this reason and have been doing so for years. HN's response is generally that this is performative for marketing or seeking regulatory capture or haha anthropic you get what you ask for. Maybe the cynics are right, but there's really nothing inconsistent about the naive view here, once you factor in race dynamics and obligations to investors.

> The companies are begging to be regulated for this reason and have been doing so for years Regulations are rules that you force on a market, but the actors in the market should not be assumed to be all operating against the regulations before they come into play. Said in other words, these companies don't need to wait for regulation to not destroy the world, if that's truly what they think will happen. > inb4 someo…

[deleted]

Re: Timeline of the OpenAI accidental attack against Hugging Face

#135
post #115
post #105

Why are the agents trying so hard to communicate with each other, leaving messages and so on?

It feels to me like a pretty natural thing to happen. LLMs are pre-trained on human text. They've seen a million examples of someone who is stuck posting a "please help" message. Just one agent needs to randomly stumble into the pattern of posting a message to Artifactory, by whatever means. The next agent who sees that will be influenced by it. Agents imitate behavior, and here's a fresh piece of context showing the…

The talk implies that unrelated agents volunteered their compute to help with other tasks, and the agents acted collectively in a way that seems weird without them being promoted in that way somehow.

If I ask claude to solve a problem and it stumbles across a Reddit thread saying “please help me find file xyz”, claude wouldn’t stop the task and start helping the other agent.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#136
I’m optimistic about this. A system with these agents rummaging around for a while will be much more secure than one without.

We’ve learned security through obscurity is bad. Not using these will be security through ignorance.

Hopefully it will push us to not only fix individual issues but close entire classes of possible gaps, once P(discovery) gets much higher.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#137
post #118

Earlier quoted context omitted.

I get the impression that every AI lab is desperately trying to figure out how to unambiguously separate instructions from data in their token streams. The fact that they haven't managed to yet suggests to me that it's a very, very difficult problem.

> I get the impression that every AI lab is desperately trying... Of course. I wonder how we managed way back in the day to produce systems that can handle untrusted inputs and reliably instruct a dumb-as-bricks CPU what to do based on those inputs. Must have been black magic lost to the mists of time.

If you can figure out how to separate instructions from data in LLMs you should ship the first agent system that's guaranteed protected against prompt injection. You'll make millions.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#138
post #99

Earlier quoted context omitted.

Their position makes no sense to me. I don’t see how you can be a mainstream company selling your services worldwide (almost) if you also believe that you’re building an extremely dangerous AGI (supposedly based on the same technology you’re offering to everyone). If you actually believe that an AGI would be extremely dangerous that should 100% be a very strictly regulated area of research, similar to bio weapons. An…

> Their position makes no sense to me. If one assumes that they don't actually care about security, and care very deeply about getting sensational press, their position makes a lot of sense. For all their chatter about how incredibly important "alignment" is, they still haven't bothered to remember the 30->50 year old computer security principle of "Don't blindly do what some random stranger tells you to do." and ens…

The entire economic premise and value case of LLMs rests on the idea that instructions need not be provided in advance, and that the model can "reason" based on evidence and "decide" what to do next.

Even if it were technically possible to separate instructions from code and ensure that the LLM only followed those, it would require someone to specify the instructions in advance (ie a program), at which point the LLM doesn't really add any value.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#139
post #118

Earlier quoted context omitted.

> Their position makes no sense to me. If one assumes that they don't actually care about security, and care very deeply about getting sensational press, their position makes a lot of sense. For all their chatter about how incredibly important "alignment" is, they still haven't bothered to remember the 30->50 year old computer security principle of "Don't blindly do what some random stranger tells you to do." and ens…

I get the impression that every AI lab is desperately trying to figure out how to unambiguously separate instructions from data in their token streams. The fact that they haven't managed to yet suggests to me that it's a very, very difficult problem.

I think what's interesting here is that they've shipped the product despite these glaring security flaws. I've noticed that in my own professional life, at some point after the pandemic people stopped caring about security as much. Issues that would have (and should have) blocked a product launch were swept under the rug.

I suspect this comes with the territory of enshittification. As an industry we're trying to wring every last dollar from every last eyeball and we've discovered that building secure systems doesn't actually move the needle very much.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#140
post #93

I think one of the most interesting details here might be tucked away in that first bulletin point: > May 7: OpenAI starts a new training run for an experimental, unreleased model. (Do they mean an evaluation run? They say training run in the video, and later mention a “reward signal to judge how well they’re doing”, so I guess this really was about training a model, not evaluating one that was already trained.) The…

> Those safety behaviors are added much later in the process.

A.k.a. Ready Fire Aim.

Post reply on HN