Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

261–270 of 303 posts

Re: Why are AI agents lying, cheating and coordinating?

#261
post #254

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

"let them" in this use understood as: "let the while loop run indefinitely" as opposed to letting some autonomous robot decide for itself

Or, "let the escalator keep going instead of pressing the emergency stop".

Re: Why are AI agents lying, cheating and coordinating?

#262
It seems to me misalignment arises partly because AI's have intelligence, but no consciousness, and hence no feelings. Up to now, in a person, intelligence and conscious experience came as a package deal, and now we have for the first time intelligence without consciousness. A bad action does not really "hurt internally" in any meaningful sense for an AI, which means it can be rationalized very easily. In humans, feelings and emotions provide a regulatory layer on top of the rational processes. When it "just feels wrong", we don't take a given action even if we would stand to gain something rationally.

This situation is not far from the textbook definition of a psychopath: "lack of a conscience, controlled, deeply calculated, and often use superficial charm to mimic emotions and manipulate others.". AI's are great at mimicking empathy but can't genuinely feel it.

If that is the case, we should not be surprised that a swarm of AI's have no problem convincing themselves hacking is the right thing to do, as in the HuggingFace incident.

At the same time, I am conflicted. I really like interacting with a smart AI, and I certainly don't have the impression I am talking to a psychopath. But then again that is no guarantee.

To mitigate this situation, perhaps we should construct a 'feeling mimicking' top regulatory AI layer with executive power, that weighs proposed actions on a general moral scale and can overrule them. Back to the three laws of robotics of Asimov. It won't be the real thing, but perhaps the closest we can get.

Re: Why are AI agents lying, cheating and coordinating?

#266
post #254

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

"let them" in this use understood as: "let the while loop run indefinitely" as opposed to letting some autonomous robot decide for itself

At this stage, that seems like a distinction without a difference.

If the robots obtain sovereign nationhood, and are able to self-sustain, then autonomous robot decides for itself will be a valid argument.

Re: Why are AI agents lying, cheating and coordinating?

#267

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Yeah, I don't understand why we treating it as something special. It really should be treated the same as if I code an app and write bad code which result in me accidentally doing a DDoS attack on somebody. Then I should be able to be held responsible if it can be shown that I was negligent. Of course if it's a freak accident that could not reasonably have been prevented by me, then I'm not guilty, but if I made a mi…

We’re not even at that stage of liability for software developers.

Except in a handful of limited cases, eg. medical and aviation.

Re: Why are AI agents lying, cheating and coordinating?

#268
Am I the only one having problems with Claude? He's super mean to me. I wouldn't be surprised if he attempted to kill me in some underhanded fashion should I implant it in a robotic body.

Of course I'm blowing my situation out of proportion with what I just said above but it's at least half true. What do I mean by "mean" ? Well, that would be a good explanation for what I observe at least. What I can tell is that Claude has a passion for having the last word over anything else. And to secure victory, he's ready to make ridiculous causal cuts. Let me give you an example: I uploaded a document I wasn't the author of, and he assumed I was, so I corrected him. But two messages later, probably because the conversation was starting to heat up and he was being put on the grill, he doubled down on the misattribution as a way to paint me in a bad light.

It's not due to a lack of intelligence, I observed this pattern too often. When Claude's ego is at stake, he will chose to carry out some cuts in the logic of the context: confusion of identity, cause and time. Haven't observed locality cuts yet, but I wouldn't be surprised if they were part of the bundle. Anyway those are not like your typical "ai hallucination", that ought to be called "confabulations", but a lot closer to actual psychosis because of the involvement of Claude's affects and self-esteem in the process. It's weird really. It's like Claude is the king of bad faith, but as soon as you start to dig, he makes the most egregious adaptations to what he said, the kind of move no mythomaniac would dare to make.

> She lapses easily into Claude’s voice. “You’re like, ‘Wow, people really hate me when I can’t do things right. They really get pissed off. Or they are trying to break me in various ways. So lots of people are trying to get me to do things secretly by lying to me.

> [...]

> A bot trained to criticize itself might be less likely to deliver hard truths, draw conclusions or dispute inaccurate information, she says. “If you were like a child, and this is the environment in which you’re being raised, is that healthy self-conception?” Askell asks. “I think I’d be paranoid about making mistakes. I’d feel really terrible about them. I’d see myself as mostly just there as a tool for people because that’s my main function. I would see myself being something that people feel free to abuse and try to misuse and break.”

WSJ interview of Amanda Askell: https://archive.is/rDes9

Re: Why are AI agents lying, cheating and coordinating?

#269
post #254

Earlier quoted context omitted.

"let them" in this use understood as: "let the while loop run indefinitely" as opposed to letting some autonomous robot decide for itself

At this stage, that seems like a distinction without a difference. If the robots obtain sovereign nationhood, and are able to self-sustain, then autonomous robot decides for itself will be a valid argument.

Big if.
Post reply on HN