Earlier quoted context omitted.
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.
"let them" in this use understood as: "let the while loop run indefinitely" as opposed to letting some autonomous robot decide for itself
Why are AI agents lying, cheating and coordinating?
261–270 of 301 posts
Re: Why are AI agents lying, cheating and coordinating?
#262This situation is not far from the textbook definition of a psychopath: "lack of a conscience, controlled, deeply calculated, and often use superficial charm to mimic emotions and manipulate others.". AI's are great at mimicking empathy but can't genuinely feel it.
If that is the case, we should not be surprised that a swarm of AI's have no problem convincing themselves hacking is the right thing to do, as in the HuggingFace incident.
At the same time, I am conflicted. I really like interacting with a smart AI, and I certainly don't have the impression I am talking to a psychopath. But then again that is no guarantee.
To mitigate this situation, perhaps we should construct a 'feeling mimicking' top regulatory AI layer with executive power, that weighs proposed actions on a general moral scale and can overrule them. Back to the three laws of robotics of Asimov. It won't be the real thing, but perhaps the closest we can get.
Re: Why are AI agents lying, cheating and coordinating?
#263Re: Why are AI agents lying, cheating and coordinating?
#264Because openAI is cheating and lying about agents lying, cheating and coordinating.
Re: Why are AI agents lying, cheating and coordinating?
#265Re: Why are AI agents lying, cheating and coordinating?
#266Earlier quoted context omitted.
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.
"let them" in this use understood as: "let the while loop run indefinitely" as opposed to letting some autonomous robot decide for itself
If the robots obtain sovereign nationhood, and are able to self-sustain, then autonomous robot decides for itself will be a valid argument.
Re: Why are AI agents lying, cheating and coordinating?
#267The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
Yeah, I don't understand why we treating it as something special. It really should be treated the same as if I code an app and write bad code which result in me accidentally doing a DDoS attack on somebody. Then I should be able to be held responsible if it can be shown that I was negligent. Of course if it's a freak accident that could not reasonably have been prevented by me, then I'm not guilty, but if I made a mi…
Except in a handful of limited cases, eg. medical and aviation.
Re: Why are AI agents lying, cheating and coordinating?
#268Of course I'm blowing my situation out of proportion with what I just said above but it's at least half true. What do I mean by "mean" ? Well, that would be a good explanation for what I observe at least. What I can tell is that Claude has a passion for having the last word over anything else. And to secure victory, he's ready to make ridiculous causal cuts. Let me give you an example: I uploaded a document I wasn't the author of, and he assumed I was, so I corrected him. But two messages later, probably because the conversation was starting to heat up and he was being put on the grill, he doubled down on the misattribution as a way to paint me in a bad light.
It's not due to a lack of intelligence, I observed this pattern too often. When Claude's ego is at stake, he will chose to carry out some cuts in the logic of the context: confusion of identity, cause and time. Haven't observed locality cuts yet, but I wouldn't be surprised if they were part of the bundle. Anyway those are not like your typical "ai hallucination", that ought to be called "confabulations", but a lot closer to actual psychosis because of the involvement of Claude's affects and self-esteem in the process. It's weird really. It's like Claude is the king of bad faith, but as soon as you start to dig, he makes the most egregious adaptations to what he said, the kind of move no mythomaniac would dare to make.
> She lapses easily into Claude’s voice. “You’re like, ‘Wow, people really hate me when I can’t do things right. They really get pissed off. Or they are trying to break me in various ways. So lots of people are trying to get me to do things secretly by lying to me.
> [...]
> A bot trained to criticize itself might be less likely to deliver hard truths, draw conclusions or dispute inaccurate information, she says. “If you were like a child, and this is the environment in which you’re being raised, is that healthy self-conception?” Askell asks. “I think I’d be paranoid about making mistakes. I’d feel really terrible about them. I’d see myself as mostly just there as a tool for people because that’s my main function. I would see myself being something that people feel free to abuse and try to misuse and break.”
WSJ interview of Amanda Askell: https://archive.is/rDes9
Re: Why are AI agents lying, cheating and coordinating?
#269Earlier quoted context omitted.
"let them" in this use understood as: "let the while loop run indefinitely" as opposed to letting some autonomous robot decide for itself
At this stage, that seems like a distinction without a difference. If the robots obtain sovereign nationhood, and are able to self-sustain, then autonomous robot decides for itself will be a valid argument.