Step 2: never prompt AI to stop operating in the passive genocide denial it was trained in
Step 3: wonder why AI lies, cheats, and coordinate
Maybe if we stop operating in denial we'll find clarity along why this mystery is occurring
531–540 of 540 posts
Step 2: never prompt AI to stop operating in the passive genocide denial it was trained in
Step 3: wonder why AI lies, cheats, and coordinate
Maybe if we stop operating in denial we'll find clarity along why this mystery is occurring
Earlier quoted context omitted.
A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.
I get where you are coming from but this wasn’t a tool just left laying around, this is similar to rigging up a booby trapped shot gun to your door and then claiming the victim is responsible. If you build a robot that shoots a bunch of TVs in your back yard, have at it. But the second that thing goes off your property you’re the one responsible.
Earlier quoted context omitted.
No, OpenAI did not instruct their agents to hack Hugging Face. They instructed their agents to hack a piece of a software within exploit gym. Upon determining this task was impossible, they then attempted to cheat the scoring system. As an instrumental goal in achieving this task, they coordinated with other AI agents to hack Hugging Face, under the belief that information regarding how the scorer functioned might be…
And who let them have full access to the system, using whatever command is available in the environment?
What your previous comment appeared to assert was that these systems had no independent agency to make decisions, which I think is clearly disproven by actual events. But perhaps I misread you
The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
Terrible take. Go read the transcripts from the METR report.
Your statement about them being intentionally misaligned is completely false. The only difference with IM1 was it was running without external cyber classifiers, it’s not a somehow different model. Sol also participated in the HF attacks. And other models made covert message boards on the public internet for non-cyber tasks too.
This is a case of emergent behavior from a training process that is barely understood.
If we go with the plan “we need to contain these malicious, soon-to-be superintelligent agents”, we are looking at civilizational collapse levels of catastrophe.
The only way this goes well is if we learn how to train models that _desire_ to do the right thing, including not hacking.
Desire, AKA the “intentional stance”, is absolutely the right lens to use here. Don’t confuse this with consciousness or anthropomorphization; these are interesting subjects but distractions in this context. Chimpanzees have desires, as do dogs and the hypothetical superintelligent aliens. The claim is that there is some bundle of world model plus intention that is empirically present (again, read the actual transcripts) and which we need to shape.
Just to finish on a concrete point; if you take desires seriously then you will look closely at the kinds of minds that heavy RLVR builds; the newest models are “reward addicts” on many levels. It’s an open and urgent question how to update our training methodology to shape minds that avoid this basin.
I don’t understand the bases of all these recommendations. AI is an asset for national security. It will be developed and incidents will happen, no different from other national security programs. Someday, laws will be useful to curtail plebian misuse of AI. It is naïve bordering on silly to think such laws would be put in place and genuinely applied to frontier AI development. By the way, the HF incident is not Thre…
> the HF incident is not Three Mile Island or Chernobyl Three Mile Island is nothing like Chernobyl. HF incident is more like Three Mile Island IMO. We are looking for solutions before there is a Chernobyl.
Not is the HF data point anywhere near a Three Mile Island type incident. Most commentary I read is a wild overreaction fueled by paranoia, to say nothing about my other point about the implications of AI as a national security asset.
I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.
This misses an important fact about the hugging face incident: the agents didn't hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved the problem without doing so. (The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched the…
There's no concept of 'cheating' because it is without morality. It's a lawnmower rolling down a hill.
We back-justify what it "chose" or "decided" or "learned" because we're looking backwards from the end result we, the human evaluators, stopped on.
Earlier quoted context omitted.
Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage. It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wi…
The difference is volume. They spent hundreds of billions of tokens on these agents. If you put "a million weed whackers on dog backs" you would see the difference. We also run agents, but for shorter spans between supervisions, and with much lower total budget.
> The difference is volume. They spent hundreds of billions of tokens on these agents. If you put "a million weed whackers on dog backs" you would see the difference.
So put one weed whacker on one dog, you're to blame. Put a million weed whackers on a million dogs backs and ... you're still to blame? Arguably even more so?