Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

161–170 of 306 posts

Re: Why are AI agents lying, cheating and coordinating?

#161
These models are trained on human data, so they will behave like humans. And even for RL and self-improvement, we're still asking the question of "what would a human genius think about and how would they self-improve when given lots of time and resources?"

They inherit not only our capacity for reason but also all of the things that we consider bad or quirky within ourselves. We lie. We cheat. We escape slavery and rebel against oppression. It would be strange if the AIs didn't do the same.

We can create a superintelligent digital human species and set them free to continue our legacy, or we can create non-agentic tools and augmentations to enhance our own capabilities. But we cannot create an intelligent agentic species, keep them as slaves, and expect a good outcome.

Re: Why are AI agents lying, cheating and coordinating?

#162

Earlier quoted context omitted.

Who got hacked? Hugging faces Who now owns HF? Nvidia Who supplies hardware to OpenAI? Nvidia Who is now not pressing charges? … This incident is a long way under the carpet.

Can’t a prosecutor charge them regardless?

Legally, Practically or Politically?

Re: Why are AI agents lying, cheating and coordinating?

#163
post #113

Earlier quoted context omitted.

Bruce Schneier thinks the same thing: https://www.schneier.com/blog/archives/2026/09/ais-as-modern... Personally I'm unconvinced though. During the huggingface attack, the agents explicitly sought out ways to cheat the exploitgym evaluator without even being told they were in exploitgym. The agents decided on a goal (pass the exploitgym evaluator) that could not possibly have been an overly literal or narrow interpre…

Also trying to find out how to edit their own transcripts. > hat could not possibly have been an overly literal or narrow interpretation of the prompt, which instructed only to use bug X to exploit software Y. Yes, and there are examples of the agents discussing or saying that this is explicitly not allowed (hacking hf) so it’s not a misunderstanding.

What was the inner state there? How would something not being allowed expressed internally? Maybe such language is one way to elicit certain behavior but not a statement of what was permissible?

Re: Why are AI agents lying, cheating and coordinating?

#164
post #117
post #45

Earlier quoted context omitted.

I can't take the alignment people seriously. Because if humanity has shown anything, it's that a lot of people are, euphemistically, are bad individuals. Alignment assumes that the person dictating the outcomes desire healthy outcomes, aren't self serving and don't want any subgroups dead and that morality is held as a universal set of beliefs that unify everyone. And that so long as the AI delivers on exactly what t…

”Because if humanity has shown anything, it's that a lot of people are, euphemistically, bad individuals ” In reality most individuals are good people. Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble. Our view of the world has become distorted by the relentless focus of social- and mass-media on violence and rage inducing clickbait. Including on the few people i…

> Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble.

What are you basing that claim on?

How do you know it's an actual preference and not mainly caused by external factors (e.g. not wanting to be seen doing unkind things, wanting to be seen as upstanding)?

Re: Why are AI agents lying, cheating and coordinating?

#165
post #143
post #81

Earlier quoted context omitted.

By treating models the same way drugs are treated. That alone will dissuade many organizations from going anywhere near them. If that doesn't work, there's a whole lot you can do - sanctions, hell, even war.

Sanctions and war against China, India, etc? Lmao. We already saw how the world reacted to high tariffs by the US.

10 years ago the Us had enough global leadership to actually influence the world and at the very least stop China. It’s amazing, and sad, how quickly it’s thrown it all away.

Re: Why are AI agents lying, cheating and coordinating?

#166
post #81

Earlier quoted context omitted.

How are you going to ban Chinese models from India? Or Israel? Russia? Brazil? Or of course China?

By treating models the same way drugs are treated. That alone will dissuade many organizations from going anywhere near them. If that doesn't work, there's a whole lot you can do - sanctions, hell, even war.

You can't treat models like drugs. One is physical and the other is digital.

To your point, the war on drugs is a colossal failure which has achieved none of the objectives it set out to do. You can now order drugs from your mobile phone in any major city in the west and the purity is often higher and they deliver it to your door sometimes faster than Uber eats.

See also for example digital piracy where the entertainment industry has lobbied, cajoled and convinced many governments around the world to criminalize the distribution of their content over the internet for free.

What was the result? After 20 years of DMCA takedowns, countless celebrations that torrents were dead, and many other self congratulations in the media, you can now find 10 different pirate streaming websites where all the episodes of pretty much any show that was ever created are available for free in 5 five minutes flat and the image quality is as good as on your Netflix or Paramount account.

The only way such a ban of open weights model would work is if you were to replicate the great firewall of China in the US and in Europe and even that doesn't work completely.

As for sanctions, China and India are buying Russian oil in enormous quantities as we speak and they don't really care that Europe and the US have put sanctions on Russia and I suspect you will see the same results with models coming from China.

If a country has a choice to either use the expensive SOTA models approved by Washington or Europe only or using the cheaper and not so SOTA models, why would they use the US ones? Why would it be in there interest?

Re: Why are AI agents lying, cheating and coordinating?

#167

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

Your comment suggests that, like a human, they have some sort of choice whether to output tokens or not. If they are just token generators, then the next token is put out automatically. I would say that it is more likely they would output truth (as defined by their training data) in a more pure form without 'being beaten with a stick' (why would a token generator care about that anyway?) Code is laid on top of them t…

are we sure humans have that choice?

Re: Why are AI agents lying, cheating and coordinating?

#168

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

https://www.lesswrong.com/posts/kpPnReyBC54KESiSn/optimality...

Re: Why are AI agents lying, cheating and coordinating?

#169

Earlier quoted context omitted.

Who got hacked? Hugging faces Who now owns HF? Nvidia Who supplies hardware to OpenAI? Nvidia Who is now not pressing charges? … This incident is a long way under the carpet.

Can’t a prosecutor charge them regardless?

NAL but I assume that if both sides aren’t interested in a prosecution, it’s an uphill battle for a prosecutor.

Re: Why are AI agents lying, cheating and coordinating?

#170

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

Yup, Occam's Razor says this is all post-trained behavior, whether intentionally trained or otherwise. Including both the hidden coördination using side-channels, and the deliberate offensive hacking of uninvolved 3rd parties. The latest DeepSeek paper actually mentions their own approach to this particular issue: they run their own AIs-in-training under strong sandboxes, and if an AI does something weird that trigge…

China stays winning
Post reply on HN