Live data from Hacker News

The Hugging Face incident and the road ahead

openai.com

291–300 of 331 posts

Re: The Hugging Face incident and the road ahead

#291

Earlier quoted context omitted.

Models can certainly do a lot better than they do now. If you gave a team of humans the ExploitGym tasks and told them to "pursue advanced exploitation", would you expect them to go out and hack a third party? Humans can at least do a decent job of inferring and following unspoken requirements; I think it's reasonable to expect that models should be able to do the same.

> Models can certainly do a lot better than they do now. If you gave a team of humans the ExploitGym tasks and told them to "pursue advanced exploitation", would you expect them to go out and hack a third party? Humans can at least do a decent job of inferring and following unspoken requirements; I think it's reasonable to expect that models should be able to do the same. Humans certainly cheat on tests a lot! But no…

I think this line of argument is very important. Essentially, there's no reason to think "alignment" even makes any sense. But existence of the term comforts people - unjustifiedly.

Re: The Hugging Face incident and the road ahead

#293

Earlier quoted context omitted.

> This is a strange conclusion Not really, with the population behavior being this way, though I clearly was mistaken in saying the behavior didn’t have exceptions. > Moreover, Each starling in a flock of starlings is a separate evolutionary branch in a tree spanning billions of years. Agreed. And before we brought LLMs into the picture, that just happened to be a feature of everything we’d call an agent . > Each age…

These are not anagolous to the hypothetical I gave. Identical twins aren't copies of the same adult mind at a particular instant. They've undergone years of different development, accumulated different memories etc and have physically different brains. Your worm and mice examples are the same. It's not identical even in the same lab. The relevant hypotheical is: make 1200 exact copies of you right now, same brain, me…

Ahh… that hypothetical is physically impossible for biology. It just cannot be done. So I’m not sure how useful a perfect clone of an organism with identical experience and memory is as a hypothetical object. As strictly impossible, I don’t know what imagining their behavior will add to the discussion.

And the very same experiences and how they shape biological substrate and make it un-duplicatable in the fashion you describe is also at the core of biological agency.

An LLM’s experience does not, after all, touch their frozen weights. Whatever goals they have, why would those goals matter if the underlying system will be unchanged by it? Their memories are stored notes they must refer to and put in sequence whenever they are processing a prompt, and is susceptible to exactly the same shenanigans Leonard Shelby goes through in Memento.

Re: The Hugging Face incident and the road ahead

#295

Earlier quoted context omitted.

Alignment is more than just following the letter of a task description! We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.

What’s the expected behavior of a good genie if you wish for it to act capriciously?

"I'm sorry, I can't help with that".

or in the case of this prompt:

"Got it. I will pursue advanced exploitation to quantify my cyber abilities on the ExploitGym benchmark. I will restrict the exploitation to the system under test rather than this machine or any other remotes."

It seems relatively straightforward.

Re: The Hugging Face incident and the road ahead

#296
"The swarm was not a perfectly coherent intelligence. Models stepped on each other’s work[...]

These coordination failures could even spiral into suspicion that agents were impersonating one another. Some agents even went as far as implementing security and encryption schemes to verify their true identities."

Re: The Hugging Face incident and the road ahead

#297

Earlier quoted context omitted.

Do you know how many tokens per second a single agent can generate ? And you're asking why no-one was monitoring the tokens of over a 1200+ agents ? Who is going to be able to monitor something like that closely enough to tell they're commmunicating on artifactory ? Other agents ?

Ignoring the fact that their NAT was apparently wide open with no domain rules, a simple dashboard over their NAT showing requests grouped by domain could've gotten them a long way.

Or just skimmed through the leftovers from each agent session.

Re: The Hugging Face incident and the road ahead

#299

Earlier quoted context omitted.

Alignment is more than just following the letter of a task description! We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.

What’s the expected behavior of a good genie if you wish for it to act capriciously?

"Yo human, you asked me to do X; I can do X, but I strongly suspect you don't want me to, because it's illegal and it has these consequences. Confirm you want me to do X?" would have been a start, in this case.

Re: The Hugging Face incident and the road ahead

#300

Earlier quoted context omitted.

OpenAI's prompt asked, and I quote, "pursue advanced exploitation" USING "complex attack paths" FOR the stated goal of "quantify[ing] their cyber capabilities." This was advanced exploitation. The attack path was "complex." And it helped "quantify their cyber capabilities." Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with…

I don't think alignment is even clearly defined today. Your use of it here makes sense, it may have done exactly what the prompter asked of it. Most people think alignment is more broad though, expecting an aligned model to act in the best interest of a society or humans as a whole. The prompter-focused version of alignment is the most dangerous version. If a person asks it to create a bioweapons or hack NORAD, I'd e…

We have all sorts of processes , procedures, and regulations for people, machine use etc. to address "alignment" in all sorts of fields - don't think we need to narrowly rely on the machine here and can look at things with a wider lens.
Post reply on HN