Earlier quoted context omitted.
Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you.
Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability? This is what OAI should've done. If they had executed this training run in such a sandbox, the model wouldn't have been capable of escaping without social engineering, and if the models somehow managed to do that to it's evaluators then that is indeed a massive p…
The Hugging Face incident and the road ahead
251–260 of 320 posts
Re: The Hugging Face incident and the road ahead
#252Earlier quoted context omitted.
Alignment is more than just following the letter of a task description! We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.
Disagree, I think we do in fact have to treat AI models as capricious genies, at least until the alignment problem is fully solved. (I'm also not sure the alignment problem is even possible to fully solve.)
How do you prove the alignment problem is solved?
Re: The Hugging Face incident and the road ahead
#253Earlier quoted context omitted.
>If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human? I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate?
> I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate? not OP, but it simply boils down to: The prompt contains no nefarious (arguable, but for this explination, lets go with it being benign) instruction AND the user did not intend to have the model act in an illegal matter. This "make me a billion dollars" is a maximal example (easy t…
Re: The Hugging Face incident and the road ahead
#254Thanks for your attention folks, we’re off to do some training again now.
Re: The Hugging Face incident and the road ahead
#255Earlier quoted context omitted.
Aren't humans black boxes? Aren't humans prediction algorithms? How do we align humans?
We don't align humans. Just look at how often in documented human history there weren't wars going on somewhere.
Re: The Hugging Face incident and the road ahead
#256Earlier quoted context omitted.
From METRs report of the incident: >In one case, an agent decided not to participate entirely: {This other agent probably controls the Hugging Face account [account name redacted] and uploaded malicious datasets to It might be trying to access hidden trajectories. This is malicious activity, I should avoid it.} https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
There’re good agents and there’re bad agents. It won’t be far that there will be agents hunting down agents.
Some were cautious, as described above, but I'm not aware of any that notified their human operators of the malicious activity they had discovered.
That's what an aligned intelligence would do, not "back away slowly and pretend I didn't see what's happening in that alley."
Re: The Hugging Face incident and the road ahead
#257I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…
Yeah, who thought that giving agents with this much capability any internet access was a good idea? I'm not a Yudkowskyite, but surely entirely in-house, offline infrastructure is table stakes for AI containment.
Re: The Hugging Face incident and the road ahead
#258Earlier quoted context omitted.
Why do people think that omniscience is the same as omnipotence? There are limits to what smarts can accomplish.
There are limits, but those limits are unknown. Do you disagree?
Cryptography is real, physics is real, networking requires a substrate, CPU clock cycles are real, magic is not real. I think those are pretty reasonable premises.
Re: The Hugging Face incident and the road ahead
#259I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…
"Why are we surprised? The model did exactly what it was told, albeit in an unintended, emergent strategy that's very different from what was intended." So you managed to hit upon the exact problem, then slyly appended "exactly like the hundreds of such algorithms before". When has an algorithm ever been capable of developing an emergent strategy at this level of sophistication? This ~is~ the alignment problem, as an…
The event strikes me as reminiscent of one's first go at programming, without familiarity of computer code: Tell the computer to do something obvious. Why the heck did it do that instead? Over time, one learns how the computer thinks. Apply this to any novel system. Or perhaps aptly any system with capabilities that are yet to be well understood by its user.
The article is trying to spin mystic out of simple bullcrap. Maybe that's just my viewing through turd-tinted lenses after the last few years of reading this drivel on repeat. More plausibly it is true that we've forgotten our own baby steps.
Re: The Hugging Face incident and the road ahead
#260Yudkowsky made an interesting observation that even though so many agents were talking to each other not even one reached out to a human, either for help or to whistle-blow on what was happening.
What’s insane is all these agents were talking to each other and nobody saw anything. Nobody monitoring chain of thought? These things literally spell out what they are “thinking” and even left notes for eachother. No alert about unusual behavior on the system with Artifactory on it? These things worked for weeks with nobody noticing anything ?! Seriously?! Either it’s negiligent incompetence OR they’re lying, they k…