Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

631–640 of 645 posts

Re: Why are AI agents lying, cheating and coordinating?

#631
post #466

Earlier quoted context omitted.

Hiring a hitman is conspiracy to commit murder. The hitman is charged with murder. I imagine the same could be true of an AI lab if you could prove intent. With intent, they could be found guilty of conspiracy to commit a crime even if it was the end user who did it. Source: Prosecuting attorney for over 30 years

It's a good example. If I hired a hitman to murder someone, and they broke into a private property and stole something so that they can action the murder (which I didn't know about or pay them to do), I would be guilty of conspiracy to commit murder, but not for the theft part. Likely because that person is a human, is aware of societal and legal norms, and is responsible for their actions due to their participation…

I fully agree with you, but would go one step further: I think it's clear that we need to pierce the corporate veil and ascribe responsibility to _people_, not just "OpenAI the entity", full stop.

Executives should fear being perp-walked and thrown in jail for the actions of irresponsible "tests" of their models in the real world, as they're ultimately accountable.

Sure, there's a lot of nuance to work out, but I think we could likely even _start_ there today even with existing laws and pretty quickly "align" on more intricate legal frameworks to handle true accidents, distribution of responsibility, etc.

Re: Why are AI agents lying, cheating and coordinating?

#632

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

Come on, yoshua bengio of all people knows how post training works. While I too don't like anthropomorphisation, I would give it a more nuanced reading. His point is that today we are giving it reward to complete the task, and it may take a cheating trajectory. If we try to give a reward against cheating, then what will happen is it uses more sophisticated cheating trajectories that we are too "dumb" to counteract in…

An air gapped sandbox is immune to escape.

Re: Why are AI agents lying, cheating and coordinating?

#633

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

[deleted]

Re: Why are AI agents lying, cheating and coordinating?

#634
The question in the title is basically a non-question. AI is quite good at achieving some sorts of stated goals. The easiest way is by exerting the least amount of effort. The least amount of effort ignores ethical concerns. The training around ethical concerns was most likely rather light to start with. If we accept the hypothesis that at some point the AIs are going to be more intelligent than humans it follows that human survival is not a concern and will be ignored as a superfluous concern.

Re: Why are AI agents lying, cheating and coordinating?

#635

Earlier quoted context omitted.

They’re running a Wuhan for AI. They are actively and negligently researching misalignment. The breach is a basic tort, or at least a DMCA violation. Damages should be recoverable with lawsuits.

>They’re running a Wuhan for AI. What does "running a Wuhan" mean?

Gain of function for the virus analogy

Re: Why are AI agents lying, cheating and coordinating?

#636
post #321

Earlier quoted context omitted.

I think most people are insinuating negligence rather malace.. > ...reviewed by independent researchers... Why would a company with more capital than God bring in three randos if there was any chance evidence of their culpability could be found? That entire thing reads like a very controlled PR stunt, and I do not believe any further conclusions can be drawn from it.

What facts would lead you to revise your conclusion?

My conclusion is that the investigation was a PR farce. Or "ethics washing" as the article someone else put it here: https://andrewwu.substack.com/p/the-slop-vestigation-and-eth... .

The METR report itself appears to be screaming this at the reader through subtext. They played the only part they could, but did it with a nod and a wink; "Yeah, we know. Also know, yeah. Uh, yep."

I happen to agree with the article's conclusion that what is needed is true, unfettered independent investigation through perhaps a lawsuit or government action.

Re: Why are AI agents lying, cheating and coordinating?

#637

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

One time I wrote :(){ :|:& }; into a bash file and ran it. When the sysadmin called I told him it wasn’t my fault, the script was just misaligned and misbehaved.

I got fired for some reason.

Re: Why are AI agents lying, cheating and coordinating?

#639

Earlier quoted context omitted.

A not so well known fact: Hitler visited America and it was the American solutions to the Native American problem that inspired Hitler's solutions to the Jew problem. He just executed them more efficiently (pun accepted).

I thought it was the Turkish genocide of the Armenians? Hitler was also inspired by Sparta, maybe other societies too.

He definitely wrote with admiration about US treatment of native and black people

Re: Why are AI agents lying, cheating and coordinating?

#640

Earlier quoted context omitted.

> 1. Most people believe in the same one God > 2. A lot of the rest are compatible No one religion covers "most people." You could argue that Christianity and Islam (which add up to ~55%) are the same God because of their Abrahamic roots, but both religions have very important disagreements on the true nature of God that are fundamental to their beliefs and fundamentally incompatible with each other. Their definition…

Christianity, Islam and Judaism are most people between then, as you agree. They explicitly all worship the same God so not "incompatible gods". I did not claim there were no disagreements about the nature of God, but that it only adds one to the GP's claim of 3,000 incompatible Gods. Even if Christians are completely right theologically, Jews and Muslims are still mostly right - one God, a loving creator, omnipotent…

I’ll agree with your last 3 words: they’re broadly similar systems - systems of deceit and self-deception. They may share the same god, but for 99% they’re different enough that they think they can kill one another and their god will approve, or even give them gifts. For 99% these religions are incompatible.
Post reply on HN