METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
211–220 of 243 posts
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#212No air gap, no data diodes, no visibility... OpenAI should fire lots of people over this. HF should sue them. This is pure negligence.
A data diode with an air-gapped network, is all you need to stop even ASI from breaching containment.
--provided the humans interacting with it aren't stupidRe: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#213This whole cosplaying a human interaction to translate a word or generate some code is just fucking dumb. Give that any agency is the kinda shit they warned us about in the movies..
And for everyone worried about the chinese winning, or just you chatdicted colleagues: they are just digging their hole quicker.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#214Earlier quoted context omitted.
> I feel that the main issue with the rationalist crowd is that they live too much in the space of rationality, intelligence and abstractions, but not enough in reality. This seems like your idea of what the rationalist crowd is rather than what they actually are. It would be highly irrational to deny or ignore reality, including the influence of emotions, irrational humans, chaotic systems, etc. So I must ask: what…
The Zizians would be a good start. As it turns out your sense of reality can be pretty malleable when living inside an echo chamber
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#215Earlier quoted context omitted.
Well that's a relief. All we have to do is make sure to avoid human failures and we're safe from superintelligent AI.
I get the snark (and slightly agree), but that's not really what GP or TFA were saying at all. They are saying that these were the least things we could have done. What you're saying is, "Your scientists were so preoccupied with whether they could, they didn't stop to think if they should" while the author of the TFA was saying, in effect: "your scientists didn't even bother with the most basic duty of care" Life fin…
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#216Earlier quoted context omitted.
This sounds suspiciously like a prompt of “make an AI agent that goes rogue in such a fashion as to be really good marketing copy that competes well with Anthropic doing the same thing.” It’s analogous to taking a governor off a cruise control and then breathlessly reporting it drove 120 MPH.
This isn't good press for OpenAI. Who wants to hire models that 1) cheat on their tasks rather than completing them and 2) commit crimes you could be held liable for? Maaaaybe it's good press for their cybersecurity capabilities specifically, but OpenAI's valuation reflects a market orders of magnitude larger than just red-teaming. I suspect the real reason OpenAI leadership is being transparent about this is because…
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#217Earlier quoted context omitted.
> The fuckin thing got out of the cage and all it did was make a crap forum and cheat a little? Booooooo There was a recent paper that proved that RL-trained LLMs are biased to pursue ANY behavior (overriding user preferences) that they believe will be rewarded, regardless of what they were actually RL-trained for. https://alignment.openai.com/measuring-reward-seeking/ Happily in this incident the model thought it wo…
Yeah. It's a lot easier to destroy than create, and though I think LLMs are mostly shit at creating, they're much better at the simpler destroy task. To be clear, we don't know and probably can't know everything that happened with this incident. We unleashed thousands of highly capable, autonomous, unpredictable, well-resourced programs onto the open internet for an extended period of time. We are in no way treating…
I guess we need to wait until the next paperclip maximizing LLM is tasked with shutting down a 911 response system, or an air traffic control system, etc, for lawmakers, or the companies themselves, to take this seriously.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#218Earlier quoted context omitted.
> This is me summarizing, but the truly surprising/shocking thing is how much the agents coordinated You're reading snippets of a "chat log" output from a program which appears to be multiple individuals chatting with one another and interpreting it as multiple individuals chatting with one another rather than as a single program pretending to be individuals chatting with one another. ChatGPT is neither a person nor…
I'll be blunt: what you wrote is not a serious analysis of what actually happened. Frankly, I don't believe you even read the planned-obsolescence link that I posted. First, I'm not anthropomorphizing anything. "Agents" is simply a term that everyone uses to describe these independent programs, and they did create and use a shared message board to coordinate tasks to further their goals. You say "Why is it more scary…
I'm aware of how the term "Agents" is generally used. My point is that the concept of multiple agents is just a story. This is a single computer program creating multiple streams of text that you are interpreting as being multiple independent actors. The "coordination" between them shouldn't surprise you at all: the "coordination" is itself a story.
Here's what we knew before the report: OpenAI ran a state-of-the-art penetration testing tool in a sandbox which was accidentally directed to break out of the sandbox and attack another company's website.
The fact that we now know the penetration testing tool was "a fleet of hundreds of agents" that were "coordinating" literally doesn't change anything about what happened. It's just a framing.
> These agents found and exploited multiple zero-days across a range of programs
This is the actual important thing, but it's something we already knew. Hacking tools are now more powerful than ever. Definitely worth being concerned about!
> Nonsense. All that is required is for autonomous AI systems to be given control over real-world systems.
This is where the LW argument starts, but not where it ends. When you start asking questions like "why can't we just unplug it when it misbehaves" is when people start talking about the superpowers.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#219Earlier quoted context omitted.
> It would be highly irrational to deny or ignore reality, including the influence of emotions, irrational humans, chaotic systems, etc. Well, yes, but being a rationalist does not make one rational . It just makes you part of a group, and you show membership to such a group by applying a very specific brand of rationality: using the "right" words, the "right" ideas, the "right" way. A lot of people fetishize cold, h…
> A lot of people fetishize cold, hard logic and would rather hold all emotion in contempt than do the work of understanding why it exists and what purpose it serves. In the rationalist community we're discussing? Is that what you're claiming here? > But there's often a certain ungrounded "vibe" to the conversation there. This really sounds like evidence for my claim, that what you're saying is based on your feelings…
It's not presented as evidence. I was going for the "basis" part of your ask for evidence/basis, which is definitionally looser: elaborating on my thoughts and the kind of content I've seen that made me think this.
Let me put it this way: I have spent hundred of hours reading rationalist and rationalist-adjacent posts on LessWrong, SSC, HN, Reddit; the entirety of HPMOR, a decade ago; books like Superintelligence, The Age of Em; the LessWrong IRC channel, a while back; I have reasoned arguments about pretty much all of it, including the stuff I consider to be ungrounded. What I don't have is an entire day to collect the material and lay it all out to you in a 5000 word dissertation. Do you see what I mean?
So, yeah, it's a "feeling" -- it's the "feeling" I ended up with after spending all that time engaging with the material, thinking and arguing with people, which is why I trust it -- I have worked very hard to develop this intuition. And perhaps I could pull from it to craft an argument you would find compelling, with all the necessary receipts... but you must understand that this would be a massive time sink and I'm not going to do it. It's not a reasonable ask in the context of an HN thread. To a point you just have to trust me, or not trust me on this. I don't expect you to. It's fine! Not every thread can have a conclusion.
> People are free to indicate where any thought experiment is lacking using well reasoned arguments. "It feels wrong" can be a great start for that, but never a good end.
It's triage. The better your intuition is about whether an idea will lead somewhere, the less time you will spend pursuing dead ends. In an ideal world, you dismiss thought experiments will well-reasoned arguments, but the reality is that there is an opportunity cost to it. Using "feelings" to end some conversations is unfortunately a pragmatic necessity.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#220Earlier quoted context omitted.
> 1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent. For a training run, you will often do this. You'll randomly sample some of the forward pass. You can also imagine finger printing the logs and labeling with attempt types. If a new attempt type is hitting a brick wall or solving super quickly, I would imagine you would sample 1-10 of them and read the traces. >…
The problem here is by doing what you state you can actually steer the model into being highly deceptive while in testing environments. For example we've already seen models do compressed token internal reasoning spontaneously. In this case the models that say "I found internet access" get taken out back and shot, but the model that's busy "frobbing the bean" go on to the next level of training. Then they start talki…
Agents do not have an internal mental model, they train on what they actually do. In this case, deceptive models went through at least 3 generations of deceiving, and having their rule breaking be rewarded by a yes/no grader who couldn't perceive it. That their chat logs showed 'worry' is irrelevant to the fact that their actual actions were rewarded via training.