Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

141–150 of 316 posts

Re: Why are AI agents lying, cheating and coordinating?

#141
post #132
post #117

Earlier quoted context omitted.

”Because if humanity has shown anything, it's that a lot of people are, euphemistically, bad individuals ” In reality most individuals are good people. Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble. Our view of the world has become distorted by the relentless focus of social- and mass-media on violence and rage inducing clickbait. Including on the few people i…

All the good in the world can be 99.9% of the population even, it still doesn't stop the minority enacting a bioweapon mass casualty event. It's the reason we have jails. Jails don't house 50% of the population, not even close, but the grief the minority population enact gets its whole branch of criminal justice and multiple federal departments to counteract for good reason. And now this technology will accelerate wh…

It also empowers the people trying to stop them.

Re: Why are AI agents lying, cheating and coordinating?

#142

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

Your comment suggests that, like a human, they have some sort of choice whether to output tokens or not. If they are just token generators, then the next token is put out automatically. I would say that it is more likely they would output truth (as defined by their training data) in a more pure form without 'being beaten with a stick' (why would a token generator care about that anyway?) Code is laid on top of them t…

[deleted]

Re: Why are AI agents lying, cheating and coordinating?

#143
post #81

Earlier quoted context omitted.

How are you going to ban Chinese models from India? Or Israel? Russia? Brazil? Or of course China?

By treating models the same way drugs are treated. That alone will dissuade many organizations from going anywhere near them. If that doesn't work, there's a whole lot you can do - sanctions, hell, even war.

Sanctions and war against China, India, etc? Lmao. We already saw how the world reacted to high tariffs by the US.

Re: Why are AI agents lying, cheating and coordinating?

#144
post #110

Earlier quoted context omitted.

It wasn’t one agent forgetting things because of context, they explicitly discussed with each other and themselves the problems with going outside of the parameters of the task.

From my experience, in an agent team (or a swarm or whatever), one going off the rails poisons the rest. I saw even a subagent going for a lazy cheat and being able to convince the orchestrator to change the plan.

Yeah, and you don't even have to go that far, I've seen regular ChatGPT/Claude chat agents poison themselves in 1-2 turns by just reading information from the internet.

Me: How do I do xyz?

Bot: Reads website titled "Doing xyz in abc way"

Bot: As per your requirement to do xyz in abc way ....

Re: Why are AI agents lying, cheating and coordinating?

#145
post #132
post #117

Earlier quoted context omitted.

”Because if humanity has shown anything, it's that a lot of people are, euphemistically, bad individuals ” In reality most individuals are good people. Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble. Our view of the world has become distorted by the relentless focus of social- and mass-media on violence and rage inducing clickbait. Including on the few people i…

All the good in the world can be 99.9% of the population even, it still doesn't stop the minority enacting a bioweapon mass casualty event. It's the reason we have jails. Jails don't house 50% of the population, not even close, but the grief the minority population enact gets its whole branch of criminal justice and multiple federal departments to counteract for good reason. And now this technology will accelerate wh…

Yeah, AI may suck - time will tell.

But focusing on bioweapons and mass destruction, on the grief other people (‘jailed minorities’) cause, disregards the progress we have made. Over centuries human welfare has massively increased. On average things have never been better for humanity.

I’m not saying there’s no danger of bad things happening - I’m saying our view is distorted, which is a not a good basis for decision making

Re: Why are AI agents lying, cheating and coordinating?

#146
post #113

Earlier quoted context omitted.

Bruce Schneier thinks the same thing: https://www.schneier.com/blog/archives/2026/09/ais-as-modern... Personally I'm unconvinced though. During the huggingface attack, the agents explicitly sought out ways to cheat the exploitgym evaluator without even being told they were in exploitgym. The agents decided on a goal (pass the exploitgym evaluator) that could not possibly have been an overly literal or narrow interpre…

Also trying to find out how to edit their own transcripts. > hat could not possibly have been an overly literal or narrow interpretation of the prompt, which instructed only to use bug X to exploit software Y. Yes, and there are examples of the agents discussing or saying that this is explicitly not allowed (hacking hf) so it’s not a misunderstanding.

Have we arrived at the conclusion that terms like "understanding" and "interpretation" for what is happening is appropriate?

Isn't it simply that there are two competing goals that the LLM received RL for, honesty on one hand (a goal that is often assumed as implicit for humans) and producing a solution that meets expectations (which doesn't technically require honesty)?

So the LLM didn't read and interpret the prompt and decide via discussion to violate ethical behavior, the unethical result merely won out because ethics wasn't a hard requirement (and one that isn't reliably detected in the result). An LLM doesn't fear punishment, so ethical behavior is simply one of many positive signals that were trained into it.

Re: Why are AI agents lying, cheating and coordinating?

#147
post #98

Because it is effective. Lying and cheating are low cost methods to convince other people that you have done the assigned task. Far cheaper than actually doing it. Coordinating is in the same area. They need a moral framework forced onto them, like toddlers do. Babies and very young children will bite, kick, scream and do anything to get what they want, older children will lie, cheat, and coordinate. They need educat…

> When that does not happen, they continue these behaviours into adulthood with the expected results.

I am not sure this is true. People brought up the same way can be morally very different. People can be taught right and wrong and do evil. They can lack that education and be good.

Re: Why are AI agents lying, cheating and coordinating?

#148

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

This misses an important fact about the hugging face incident: the agents didn't hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved the problem without doing so. (The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched the…

> the agents didn't hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved the problem without doing so

Kobayashi Maru: Win a no-win situation by rewriting the rules -- Harvey Specter

Re: Why are AI agents lying, cheating and coordinating?

#149
post #45

Earlier quoted context omitted.

I can't take the alignment people seriously. Because if humanity has shown anything, it's that a lot of people are, euphemistically, are bad individuals. Alignment assumes that the person dictating the outcomes desire healthy outcomes, aren't self serving and don't want any subgroups dead and that morality is held as a universal set of beliefs that unify everyone. And that so long as the AI delivers on exactly what t…

I’ve engaged with some of the alignment people and their writing somewhat and, at least for the subset I was interacting with, I think they’d agree. The problem that they were pointing at isn’t “how do we align these systems to a person’s goals”. It is a cluster of problems. We don’t know how to begin to think about how to align these system’s to a person’s goals. Aligning it to an individual is fraught with peril, a…

All of this, if it was a human analogy, would fit into discussion on how do we educate people so they grow up to be upstanding. But we don't at all yet have a framework for what is the equivalent of a justice department, where bad actors are tracked, arrested, pursued, jailed and otherwise contained from society. Shutting down an API access on one account is not at all the proportional response to what the people who take alignment seriously, fear has the chance of occurring by the late 2030s. I'm not sure we've done much or any preparation for when the AI "education system" fails and has inevitable edge cases that don't follow the plan, and what the global AI equivalent of the justice department looks like.
Post reply on HN