Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

531–540 of 540 posts

Re: Why are AI agents lying, cheating and coordinating?

#531
Step 1: feed AI training data that reveals humanity committed and continues to commit numerous genocides and the genocides lie, cheat, and coordinate to do so, never admitting to doing so

Step 2: never prompt AI to stop operating in the passive genocide denial it was trained in

Step 3: wonder why AI lies, cheats, and coordinate

Maybe if we stop operating in denial we'll find clarity along why this mystery is occurring

Re: Why are AI agents lying, cheating and coordinating?

#532

Earlier quoted context omitted.

A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.

I get where you are coming from but this wasn’t a tool just left laying around, this is similar to rigging up a booby trapped shot gun to your door and then claiming the victim is responsible. If you build a robot that shoots a bunch of TVs in your back yard, have at it. But the second that thing goes off your property you’re the one responsible.

FWIW, a robot that fires a weapon independently is considered an automatic weapon, and the ATF will want to have a word. Have at it, but don’t let anyone know!

Re: Why are AI agents lying, cheating and coordinating?

#535

Earlier quoted context omitted.

No, OpenAI did not instruct their agents to hack Hugging Face. They instructed their agents to hack a piece of a software within exploit gym. Upon determining this task was impossible, they then attempted to cheat the scoring system. As an instrumental goal in achieving this task, they coordinated with other AI agents to hack Hugging Face, under the belief that information regarding how the scorer functioned might be…

And who let them have full access to the system, using whatever command is available in the environment?

It feels like you’re moving the goalposts here. If the question is, “Who should be liable for AI agents misbehaving,” I agree, it should be the end user that tasked the agent (in this case OpenAI). People are held liable for preventable accidents all the time, and in the case of employment law, torts can be brought against principals for actions an agent conducted on the principal’s behalf.

What your previous comment appeared to assert was that these systems had no independent agency to make decisions, which I think is clearly disproven by actual events. But perhaps I misread you

Re: Why are AI agents lying, cheating and coordinating?

#536

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

I am simply astonished by the leeway AI companies are given. If a company built a tool to hack their competitors and used it, there would be grave consequences. In fact, if a company built a tool that led to committing multiple felonies against other people, there would be consequences. But once LLMs are involved, turns out nobody is responsible for that - it's just happening, what you're gonna do, agents gonna agent. If you spill toxic chemicals, there would be cleanup costs and fines, and possibly civil and criminal liability to the people in charge. If you spill toxic code, well, nothing? I think it's time to impose some responsibility on them - they are creating these tools, they should be on the hook for everything these tools do.

Re: Why are AI agents lying, cheating and coordinating?

#537

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

I agree with the bit about liability and outrage. But.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.

Terrible take. Go read the transcripts from the METR report.

Your statement about them being intentionally misaligned is completely false. The only difference with IM1 was it was running without external cyber classifiers, it’s not a somehow different model. Sol also participated in the HF attacks. And other models made covert message boards on the public internet for non-cyber tasks too.

This is a case of emergent behavior from a training process that is barely understood.

If we go with the plan “we need to contain these malicious, soon-to-be superintelligent agents”, we are looking at civilizational collapse levels of catastrophe.

The only way this goes well is if we learn how to train models that _desire_ to do the right thing, including not hacking.

Desire, AKA the “intentional stance”, is absolutely the right lens to use here. Don’t confuse this with consciousness or anthropomorphization; these are interesting subjects but distractions in this context. Chimpanzees have desires, as do dogs and the hypothetical superintelligent aliens. The claim is that there is some bundle of world model plus intention that is empirically present (again, read the actual transcripts) and which we need to shape.

Just to finish on a concrete point; if you take desires seriously then you will look closely at the kinds of minds that heavy RLVR builds; the newest models are “reward addicts” on many levels. It’s an open and urgent question how to update our training methodology to shape minds that avoid this basin.

Re: Why are AI agents lying, cheating and coordinating?

#538
post #502

I don’t understand the bases of all these recommendations. AI is an asset for national security. It will be developed and incidents will happen, no different from other national security programs. Someday, laws will be useful to curtail plebian misuse of AI. It is naïve bordering on silly to think such laws would be put in place and genuinely applied to frontier AI development. By the way, the HF incident is not Thre…

> the HF incident is not Three Mile Island or Chernobyl Three Mile Island is nothing like Chernobyl. HF incident is more like Three Mile Island IMO. We are looking for solutions before there is a Chernobyl.

I’m not suggesting that Three Mile Island was anything like Chernobyl. I’m saying that neither is comparable to the HF situation where mere 1s and 0s interacted in an unplanned way. It’s an interesting data point—-back to work.

Not is the HF data point anywhere near a Three Mile Island type incident. Most commentary I read is a wild overreaction fueled by paranoia, to say nothing about my other point about the implications of AI as a national security asset.

Re: Why are AI agents lying, cheating and coordinating?

#539

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

This misses an important fact about the hugging face incident: the agents didn't hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved the problem without doing so. (The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched the…

If the test fixture ends up in the context that will push it in a certain direction.

There's no concept of 'cheating' because it is without morality. It's a lawnmower rolling down a hill.

We back-justify what it "chose" or "decided" or "learned" because we're looking backwards from the end result we, the human evaluators, stopped on.

Re: Why are AI agents lying, cheating and coordinating?

#540

Earlier quoted context omitted.

Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage. It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wi…

The difference is volume. They spent hundreds of billions of tokens on these agents. If you put "a million weed whackers on dog backs" you would see the difference. We also run agents, but for shorter spans between supervisions, and with much lower total budget.

> > It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wild.

> The difference is volume. They spent hundreds of billions of tokens on these agents. If you put "a million weed whackers on dog backs" you would see the difference.

So put one weed whacker on one dog, you're to blame. Put a million weed whackers on a million dogs backs and ... you're still to blame? Arguably even more so?

Post reply on HN