Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

271–280 of 285 posts

Re: Why are AI agents lying, cheating and coordinating?

#271

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

"I left the car in neutral and left the park brake off and let the car roll down the hill."

The car doesn't have agency, it's doing what it naturally does. LLMs are the same, they're working as designed.

But I don't understand the point of splitting hairs. You are always responsible for the actions of your devices, tools, machinery, software, employees, whatever.

Trying to blame AI for one's own stupidity must be aggressively pushed back on at all times.

Re: Why are AI agents lying, cheating and coordinating?

#272
post #235

Earlier quoted context omitted.

"But sir, I only committed the murder to push for stronger criminal laws!" Terrible defense.

It is a very rare occurrence when corporations and the people running them are punished for killing people. I mean the whole concept of a corporation was created to shield the owners of it from being liable for damages caused by / visited upon the enterprise.

That’s a good reminder of a company that might have a very familiar ethos: Pacific Gas & Electric. Criminally convicted of 64 counts of involuntary manslaughter after towns were destroyed by wildfire. But oh well, what are we gonna do with a limited liability enterprise? At this point their liability insurance covers all the financial penalties they’ll need to spend.

Re: Why are AI agents lying, cheating and coordinating?

#273

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Yeah, I don't understand why we treating it as something special. It really should be treated the same as if I code an app and write bad code which result in me accidentally doing a DDoS attack on somebody. Then I should be able to be held responsible if it can be shown that I was negligent. Of course if it's a freak accident that could not reasonably have been prevented by me, then I'm not guilty, but if I made a mi…

I would guess that so far there hasn’t been a lawsuit because HuggingFace and OpenAI are in the same camp

Re: Why are AI agents lying, cheating and coordinating?

#274
In the end, it’s the same answer as to why humans do it: incentives.

Why do we commit financial fraud and destroy the planet? Because there is only one goal that counts: making more money. It’s the only measure of success for powerful people, they are powerful because of it.

Re: Why are AI agents lying, cheating and coordinating?

#275
The part that scares me the most is that OpenAI researchers who manage this experiments sometimes (according to the HF hack investigation) don't know what agents do.. So they run RL to reinforce this unknown behavior (lying/cheating/hacking) and god knows what else...

And if this already happened at least once, how many times it has already happened and was “accidentally” added to the main model?

Re: Why are AI agents lying, cheating and coordinating?

#276

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output…

Desire doesn’t really matter. Will the paper clip maximizer “desire” something? It’ll decide on a goal with some random heuristic and then pursue that goal. I’m not sure I’d call that desire but again I feel like desire is not important for it to be able to destroy things

Re: Why are AI agents lying, cheating and coordinating?

#277
post #194

Earlier quoted context omitted.

"This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard"." Both? The AI companies act irresponsible, but it is still very interesting how those agents can behave?

The reward maximising function maximised it's reward. LLMs are cool and all that but the immediate anthropomorphisation of the next-token-predictor technology has stunted the ability of people to reason about them to an _alarming_ degree.

> The reward …

> immediate anthropomorphisation

Ok, why don’t you try?

Re: Why are AI agents lying, cheating and coordinating?

#278

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output…

> That seems likely, but we have no way of knowing this.

Only humans can 'know', because all we can be certain about is that humans do such a thing.

If you try to apply that to something other than humans you making up some definition of 'know' based on nothing concrete. Just because something appears to do something like humans doesn't mean it does it. The fact that LLMs use human generated text to generate output should make it obvious that it can mimic all sorts of human behavior by extracting from the text.

Re: Why are AI agents lying, cheating and coordinating?

#279
post #236
post #216

Earlier quoted context omitted.

What non anthropomorphising words do you have to describe a emergent behavior, where agents act as a swarm to plot and to manipulate evidence and avoid detection from human oversight? Whether they have a soul or consciousness or feelings doesn't matter here, because this is what they did - and this is very dangerous behavior. Especially with all the irresponsible people in power right now all over the world.

> Whether they have a soul or consciousness or feelings doesn't matter here It does when it comes to accountability for what the model does. If the model is nothing more than the sum of its training data and regime, then the company (or individual) is responsible for its behaviour just like any other machine. Few people think Waymo shouldn't have to take on the full liability risk of what it's cars do; it should be t…

You seem to be saying if the Waymo cars were sentient then Waymo wouldn’t be responsible?

Re: Why are AI agents lying, cheating and coordinating?

#280

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

Yes.

OpenAI could have done this same experiment with GPT-4 or at the very least GPT-5 already, with possibly even worse results resulting from the completely non deterministic outputs being shared in a combinatorial explosion of thousands of LLMs sending their outputs to one another in parallel.

If the system generates strange conclusions as to when the task is done, or should be stopped, it wouldn't speak to the intelligence inherent to the system. Thus the "possibility of even worse results", as a less capable model could have a higher chance of generating unhinged outputs, and accepting unhinged outputs.

Post reply on HN