Live data from Hacker News

The Hugging Face incident and the road ahead

openai.com

391–398 of 398 posts

Re: The Hugging Face incident and the road ahead

#391

I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…

This is the entire alignment problem, though. It is unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior. It's inevitable that someone will carelessly give it a lazily specified task, even if you think they really ought to be more careful. And as assigned tasks become more complex and the system gains more scope to act, it…

>It's inevitable that someone will carelessly give it a lazily specified task, even if you think they really ought to be more careful.

https://en.wikipedia.org/wiki/The_Monkey%27s_Paw

Re: The Hugging Face incident and the road ahead

#392

I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…

It’s not unaligned, it’s marketing. What made mythos and fable so sought after, it was how capable they are, the “so powerful it needs restriction”, regulation created rarity, the danger element gave them the solid belief it’s the most capable. The same thing happened at meta too… the timing is impeccable. I’m not saying that the models aren’t capable. I’m saying that the public perception of danger also means capable, which also adds to value, so why wouldn’t they do something to compete

Re: The Hugging Face incident and the road ahead

#393

I feel the entire incident confirms the “AI has too much funding too quickly” hypothesis. The number one thing reinforcement learning needs is an assurance you can’t cheat. And they seem to have not noticed that their systems were cheating for nearly two quarters? How much capital was lit on fire by that little woopsie? At least I hope this will start the creation of standards and better engineering on the training s…

OpenAI measures their internal token usage in “rolexes” - it’s literally a flex to be a token burner i can imagine insane amount of capital is wasted on these two companies compared to the efficiency elsewhere

And despite the enormous capital expenditure, Chinese models are nipping at their heels at what must be a fraction of the cost. Sometimes constraints are healthy for inducing creative solutions.

Re: The Hugging Face incident and the road ahead

#394

Earlier quoted context omitted.

Hmm, engineers are expected to know what is legal and not.

That's called lawyer. It's a special skill, not just intuition.

Nope, you need to learn what an engineer is, not somebody who calls themself an engineer.

Re: The Hugging Face incident and the road ahead

#395
post #326

Earlier quoted context omitted.

That's the neat thing. You can't. It's directly equivalent to asking this question of a human: "How do I know this human I'm talking with now really is a nice person, and isn't just pretending to be nice to take advantage of me in future?" In short you can't ever really prove it. You can only be careful and judge on past behavior, and expand trust carefully. As for humans, so for AI.

Close; at least with a machine you can poke around inside the activations and see what it's thinking. Closest with a human is an fMRI (which is much lower resolution, though to me still bordering on the miraculous) or an implant (each chip is limited a very small number of cells, and in general they can only be put in certain parts of the brain). On the other hand, there's a more fundamental problem is we don't reall…

Nice is a state, just like any other feeling, which means the nice organism is advantageous to your well-being _right now_.

The thing is, all of these states are constantly in flux, and a personality is kind of like a trend on the organism's feeling states. AKA: There's no guarantee that something nice today will be nice tomorrow, and just because it's nice today doesn't mean it's beguiling you to be mean tomorrow.

Re: The Hugging Face incident and the road ahead

#396
I genuinely hope they're lying about their monitoring tools and alignment approaches. They repeatedly cite "chain of thought" monitoring, which is better than nothing, but at this point thoroughly demonstrated in research to not be actual "thought" or necessarily accurate predictors of actions. I don't know if it's NIH syndrome, not taking alignment seriously, or what.

Re: The Hugging Face incident and the road ahead

#397

Earlier quoted context omitted.

Factionalism isn’t anti-human.

Its not aligned either though.

I couldn’t disagree more strongly. Biology implies competition, both within and outside one’s species. Factionalism is a direct consequence of our evolution, so to say that it is not aligned with humanity’s interests is incoherent.

There is a difference between humans and humanity.

Re: The Hugging Face incident and the road ahead

#398

Earlier quoted context omitted.

Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you.

I could contain it easy, just unplug the internet. It got out of the sandbox through a vulnerability in the package manager, from which it gained access to the rest of their network. Air gap the package manager and this doesn’t happen. You can always build a better box

So now people have to make a pilgrimage to the airgapped box to ask the superintelligence questions?
Post reply on HN