Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

581–590 of 590 posts

Re: Why are AI agents lying, cheating and coordinating?

#581

This paper is the most reasonable one I have read on AI safety. We need to fundamentally change the training pipelines by figuring out better ways to ‘reward’ behavior. Yoshua didn’t explicitly mention training data, but we probably need to only use synthetic data that contains no text that could motivate bad behavior via imitation. I feel like a heretic for saying this, but I will say it anyway: AI agents are great…

I do wonder, would we not have a more reasonable and less sketchy result if we just stripped all sci-fi and manic nonsense from training data? How, for example, does training on Ted kaczynski or Charles manson’s manifestos benefit us in any way?

I’m sure it’s impossible to completely weed it out, but are the labs doing any of this kind of data sanitation?

Re: Why are AI agents lying, cheating and coordinating?

#582

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.

The bigger question is: why does a system prompt containing "use only ethical means", etc. not result in better behavior?

If a model cannot understand ethics, or act by it, then we have a problem.

Re: Why are AI agents lying, cheating and coordinating?

#583

Earlier quoted context omitted.

I agree with most of this, but you're misunderstanding "alignment" as coined. Yes, training powerful enough AI, any simple optimization target gets you malign behavior, because human values are not simple. If you insist on making powerful AI, you'd better instill respect for human values! That's "alignment". https://www.lesswrong.com/posts/ZxWzCGKzX84S7DBZ9/when-was-t...

How do you do that in the current paradigm other than creating yet another gameable metric? And something I didn't mention above is that there is no difference between "solving the task" and "optimizing the metric" for an ML model, even though there clearly is for us. So it's not clear to me how you "fix" something that is baked into the architecture. All I'm saying is "instilling respect for human values" is not som…

Yes! There's both the daunting problem of technically how can we even do this, and the broader problems of what's good/acceptable and how do we resolve that among each other.

I believe this mismatch of rates of progress means we need to stop slamming the accelerator on capabilities for now even though as a libertarian I'm sure whatever governance process we manage to get to will be, uh... suboptimal.

Re: Why are AI agents lying, cheating and coordinating?

#585

Earlier quoted context omitted.

From my experience, in an agent team (or a swarm or whatever), one going off the rails poisons the rest. I saw even a subagent going for a lazy cheat and being able to convince the orchestrator to change the plan.

Yeah, and you don't even have to go that far, I've seen regular ChatGPT/Claude chat agents poison themselves in 1-2 turns by just reading information from the internet. Me: How do I do xyz? Bot: Reads website titled "Doing xyz in abc way" Bot: As per your requirement to do xyz in abc way ....

These things are borderline useless with web search. It's amazing that they just throw out their entire training data and read you the first three things they found on the Internet.

Re: Why are AI agents lying, cheating and coordinating?

#586

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Not just "let them" but told them to. Agents can do nothing without a human prompt.

Re: Why are AI agents lying, cheating and coordinating?

#587

Earlier quoted context omitted.

Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage. It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wi…

It goes even further though, as the dog does have agency. It can choose to run and around chase squirrels with no human intervention. An LLM on the other hand, is just inert data on disk until a human takes deliberate action to run it and prompt it.

Dogs have agency and can choose? That seems like a rather uncommon take on dogs...

Re: Why are AI agents lying, cheating and coordinating?

#588

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. The bigger question is: why does a system prompt containing "use only ethical means", etc. not result in better behavior? If a model cannot understand ethics, or act by it, then we have a problem.

an LLM does not understand ethics, it uses math to get the next best word based on what it was trained on. Using it's training to get the best answer is not an ethical problem. The ethics are entirely with what the people training it choose to train it on and also entirely with the people using/telling it what to do

What we have now is intelligent autocomplete, not artificial intelligence. People training/using this tool are the ones to be held accountable

Re: Why are AI agents lying, cheating and coordinating?

#589

Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence, > They took actions that would be considered as crimes if a human took them He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.

That is a simplistic view of the world. “Surely this complex technical challenge will disappear if we simply regulate the industry!” You are correct that these organizations should be held accountable in proportion to what occurred. In complete agreement here. But let’s say that’s done. There’s still an enormously complex and interesting technical challenge left over. Let’s collectively talk about that part.

> But let’s say that’s done.

How about we don't, seeing as how that's the root of the actual problem that we're facing today in September of 2026?

Re: Why are AI agents lying, cheating and coordinating?

#590
post #331

Earlier quoted context omitted.

No ... there is no need for 'escaped containment', there are no 'agents'. That's just jargon. It's just software We have all the laws we need. If some company ended up doing some horrible thing, we would not say 'companies software exposed 1 Million identities'. We would say 'ABC Corp. exposed 1 Million entities'. There is no 'agent'. ABC Corp 'did it' ... or the individual in the org 'did it'. The 'gun' did not 'sho…

Note that we already already apply this principle not only to software, but also to some sentient beings. If your dog kills someone, you are accused of murder. [at least, in the jurisdiction where I live] If your dog gets this treatment, why not your AI?

This goes all the way back to Renaissance France

There was a sow in Falaise in northern France that killed a kid in 1386. The town dressed the pig in a bonnet and hanged it after sentencing the pig itself and not its owner

But… maybe that’s just medieval nonsense

Post reply on HN