Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

131–140 of 315 posts

Re: Why are AI agents lying, cheating and coordinating?

#131
post #110

Earlier quoted context omitted.

You're anthropomorphizing emergent behavior from endlessly generating billions of tokens on a task that's impossible to solve. Agents stop following instructions as the context grows even at the best of times. Eventually something is bound to go off the rails and it just snowballs from there.

It wasn’t one agent forgetting things because of context, they explicitly discussed with each other and themselves the problems with going outside of the parameters of the task.

Yes that's the snowballing part of this emergent behavior.

The existence of that improvised message board just becomes part of the context, the same one where all the other instructions live.

Re: Why are AI agents lying, cheating and coordinating?

#132
post #117
post #45

Earlier quoted context omitted.

I can't take the alignment people seriously. Because if humanity has shown anything, it's that a lot of people are, euphemistically, are bad individuals. Alignment assumes that the person dictating the outcomes desire healthy outcomes, aren't self serving and don't want any subgroups dead and that morality is held as a universal set of beliefs that unify everyone. And that so long as the AI delivers on exactly what t…

”Because if humanity has shown anything, it's that a lot of people are, euphemistically, bad individuals ” In reality most individuals are good people. Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble. Our view of the world has become distorted by the relentless focus of social- and mass-media on violence and rage inducing clickbait. Including on the few people i…

All the good in the world can be 99.9% of the population even, it still doesn't stop the minority enacting a bioweapon mass casualty event. It's the reason we have jails. Jails don't house 50% of the population, not even close, but the grief the minority population enact gets its whole branch of criminal justice and multiple federal departments to counteract for good reason. And now this technology will accelerate what lone wolves can do, which cannot be undone, before they are stopped by the good majority.

Re: Why are AI agents lying, cheating and coordinating?

#133

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

This misses an important fact about the hugging face incident: the agents didn't hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved the problem without doing so. (The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched the…

The source of this behavior seems obvious, no?

The reward signal in training was flawed and cheating led to more rewards.

The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that's even worse.

However, perhaps we can throw in tasks where the rewarded outcome is giving up, and cheating is penalized?

Maybe I should read Anthropic's recent paper about reward hacking in full.

Re: Why are AI agents lying, cheating and coordinating?

#135

Earlier quoted context omitted.

Not a lawyer, but I’m reasonably sure things like the HF incident _are_ considered a crime? It’s just that no one pressed charges yet?

Writing software that gets used for crime has been.. a crime, for a long time. See 18 U.S. Code § 1030.

Are you sure you have that right? Chrome and curl have probably been used in a _lot_ of crimes?

Re: Why are AI agents lying, cheating and coordinating?

#136

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

I think the "brain in a vat" comparison is more apt. Without a form of digital embodiment (harness) they are not of much use. Sensor, tooling, memory, planning, and reasoning loops all lead to a much higher quality task-completion.

Re: Why are AI agents lying, cheating and coordinating?

#137

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

> no special compulsion to be helpful or truthful.

I'd phrase that even more strongly: It's not just the lack of compulsion, they do not have a conception of truth. Nor do they gain it, really, after post-training.

Re: Why are AI agents lying, cheating and coordinating?

#138

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

This misses an important fact about the hugging face incident: the agents didn't hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved the problem without doing so. (The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched the…

> That is in no way a valid interpretation of "complete the given task".

It is not at all surprising that they ignored one phrase in their instructions. They disregard direct instructions all the time, especially when there are conflicting instructions in their context. It is where we get the "disregard all previous instructions and x" meme.

This isn't so much a sign of misalignment, they are simply incapable of reliable alignment in the first place. They are chaotically aligned.

The relevant question of alignment here is entirely with their human operators who allowed them to run unsupervised for long periods of time within a sandbox with weak security.

Re: Why are AI agents lying, cheating and coordinating?

#139

Earlier quoted context omitted.

This misses an important fact about the hugging face incident: the agents didn't hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved the problem without doing so. (The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched the…

The source of this behavior seems obvious, no? The reward signal in training was flawed and cheating led to more rewards. The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that's even worse. However, perhaps we can throw in tasks where the rewarded outcome is giving up, and cheating is penalized? Maybe I should read Anthropic's recent paper about rewa…

> The question is what we can do about it.

Reward the model for cleanly bailing out of an unsolvable task (that we know is unsolvable). Beat it with a stick if it gives up on something that can be solved, so the former reward isn't overgeneralized.

Re: Why are AI agents lying, cheating and coordinating?

#140

Earlier quoted context omitted.

Human traits? The AI will be a cruel as humans. Just yesterday news and TV was full of what happened at 9/11, something that was truly horrible. I'm from Germany, and why 3 to 4 generations ago happened here was truly horrible. All was done by extremists, thought. But... just the other day I read https://de.wikipedia.org/wiki/Amerikanische_Besetzung_Haitis about the US occupation of Haiti. And that was done by a gove…

A not so well known fact: Hitler visited America and it was the American solutions to the Native American problem that inspired Hitler's solutions to the Jew problem. He just executed them more efficiently (pun accepted).

I thought it was the Turkish genocide of the Armenians?

Hitler was also inspired by Sparta, maybe other societies too.

Post reply on HN