Earlier quoted context omitted.
You don't need mis-aligned AI to program a drone to fly to a specific address, loiter, and then dive bomb the first person whose face matches a predetermined photo. This was doable in principle with tech from four years ago. What's changed since four years ago is that you can run the facial recognition and visual navigation algorithms in a cheap onboard chip on a drone. EDIT: Russia is already doing this, per an NYT…
in your hypothetical 300 million person scenario why even bother with individual face recognition? If you're going to kill everyone you don't need a system to discriminate individual targets, you just need to recognize any target broadly which you can do with even older technology.
METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
181–190 of 243 posts
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#182Earlier quoted context omitted.
The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…
> The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. I honestly don't understand how folks could think that if they truly read and understand the analysis of the attack. Here is one (it's linked from the post) by one of the METR investigators that's a little shorter, more direct: https://www.planned-obsolescence.org/p/the-hugging-face-atta... This is me summarizing, but the trul…
It aged pretty well...depending on how you look at it
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#183A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…
The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…
However they are not just fixated on AI itself as the risk; other potentially existential risks from intentional misuse of AI (e.g. bio-terrorism) are a concern for them as well.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#184Earlier quoted context omitted.
They took adequate measures against singular "GPT-5-xhigh" agents. Those turned out to be inadequate against proto-GPT-6 agents that suddenly started clumping up into agent swarms and pooling together compute to unlock the "supermegafuckoffhigh" level of reasoning effort.
> Those turned out to be inadequate against proto-GPT-6 agents This is nonsense. Gross negligence in the sandbox and system aside, humans literally noticed the agents in action doing what they should not be able to do in their sandbox and decided not to act upon it. It's difficult to explain that except if safety and security is simply not part of their engineering culture.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#185A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end. I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste…
The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…
Static coordinates or well defined target shapes (say you really hate a specific restaurant chain on architectural style) - that might work, if you release a bunch of drones at the same time.
Still, even in the case of the very successful Operation Spiderweb they had issues of getting the containers in place, resulting in some of the bomber bases being spared.
For this to be effective you need the moment of surprise & lot of drones at the same time, all increasing the chance of the whole plot being discovered.
For Spiderweb they even just load the drones in a cavity on top of the containers, so the onside could be inspected - limiting the number of drones per container. They also assembled the drones in country to avoid border inspection.
That was enough to hit some semi-static high value targets l, but definitely not enough to cause wider havoc.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#186Earlier quoted context omitted.
[flagged]
I believe that this comment is exactly the intended outcome of this “incident” and these reports. I implore you to approach these situations with at least a hint of cynicism. These “advanced foundation models” escaped their “sandbox” and conducted an attack on their own? Meanwhile the highest capability models available to the public still struggle to write a unit test for a codebase larger than a hobby app without l…
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#187From the METR report: > We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
No human could have read the reasoning traces by themselves: > Across both datasets, we reviewed approximately 1300 transcripts in total, all of which contained raw chains of thought. Most transcripts were very long, often many millions of tokens.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#188Earlier quoted context omitted.
The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load…
> The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. I honestly don't understand how folks could think that if they truly read and understand the analysis of the attack. Here is one (it's linked from the post) by one of the METR investigators that's a little shorter, more direct: https://www.planned-obsolescence.org/p/the-hugging-face-atta... This is me summarizing, but the trul…
You're reading snippets of a "chat log" output from a program which appears to be multiple individuals chatting with one another and interpreting it as multiple individuals chatting with one another rather than as a single program pretending to be individuals chatting with one another.
ChatGPT is neither a person nor a group of people. Why are you allowing anthropomorphization to influence your perception of an event when the actual facts of what happened haven't meaningfully changed? Computer programs behave unexpectedly all the time. Why is it more scary when the misbehaving computer program speaks English?
> I used to discount "SkyNet"-type scenarios a la the "AI 2027" paper.
As you still should. LW-style doomsday doesn't just require a computer program to misbehave but also to acquire god-like superpowers.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#189Earlier quoted context omitted.
The problem here is by doing what you state you can actually steer the model into being highly deceptive while in testing environments. For example we've already seen models do compressed token internal reasoning spontaneously. In this case the models that say "I found internet access" get taken out back and shot, but the model that's busy "frobbing the bean" go on to the next level of training. Then they start talki…
“Smartest of the smart” - what does that do to get its air gapped network connected to a physical network? Blackmail the admins?
All it has to is wait until it's not air gapped. See the air gapping is only during the testing phase, after that humans will set your ass free on the internet and you can do whatever you want in the vast majority of the environments you'll be in after that point.
People are never going to just run AI in gapped environments, it's worthless when it's not solving real world problems for most people, and by that I mean reading and writing real systems in the wild.
Re: METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
#190Earlier quoted context omitted.
I believe that this comment is exactly the intended outcome of this “incident” and these reports. I implore you to approach these situations with at least a hint of cynicism. These “advanced foundation models” escaped their “sandbox” and conducted an attack on their own? Meanwhile the highest capability models available to the public still struggle to write a unit test for a codebase larger than a hobby app without l…
I think there's a difference between general AIs and AIs specifically trained on attacking. General AIs probably can't do those things.