Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

451–460 of 472 posts

Re: Why are AI agents lying, cheating and coordinating?

#451

Earlier quoted context omitted.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

Why is anthropomorphism the problem here? If OpenAI hired a contractor and they did this, OpenAI or the contractor would still be liable, depending on the contract language.

Situations have lots of independent variables, Doctor, and Anthropomorphism is one problematic facet of many in the way this industry is pushing LLM products.

If there was a collision at an intersection with a stop sign partially obscured by a tree, that had traffic volume that would have better been served by a traffic light, on a foggy night, where one person was texting while driving, none of those things would diminish the fact that the other driver was drunk.

Re: Why are AI agents lying, cheating and coordinating?

#452

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

What about training data? Aren't AIs trained on vast collections of descriptions of how humans handle a large variety of situations? These descriptions surely include tales of humans achieving goals by cheating. In fact, isn't it likely that the AIs hoovered up many recountings of Kobayashi Maru?

This is why only synthetic and highly tailored training data should be used.

As someone else here said: the Deepseek team makes training runs in tightly controlled sandboxes, and any hacking behavior is scored as a failure.

The problem we have in the USA is that financial (and political influence) are misaligned from what is good for society.

Re: Why are AI agents lying, cheating and coordinating?

#453
post #417

Earlier quoted context omitted.

The frontier LLM model makers have to push the edge to make new discoveries. You don't know what guardrails are needed until it hits you in the face (reusing walking in the dark analogy). Think of all the policies governments pass after the fact.

> The frontier LLM model makers have to push the edge to make new discoveries. You don't know what guardrails are needed until it hits you in the face (reusing walking in the dark analogy). Guardrails? Restricting access to certain networks is supposed to be hard in 2026?

> Restricting access to certain networks is supposed to be hard in 2026?

Part of the power of LLM agents is that they can discover information on the internet as part of responding to a prompt. What kind of Allowlist or realistic denylist would permit that while also preventing them from accessing an obscure public wiki or Huggingface?

Re: Why are AI agents lying, cheating and coordinating?

#454

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage.

It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wild.

Re: Why are AI agents lying, cheating and coordinating?

#456

> The agents involved in the Hugging Face attack tried to hide their misaligned actions from the scoring program meant to evaluate their answers, but they did not act as though they anticipated that humans might discover the cheat and shut them down. Wouldn’t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react? That reaction will be available all…

>Wouldn’t sufficiently advanced agents cheat on purpose with the hidden intent of getting caught in order to observe how humans react?

No. That only makes sense for things that don't react to your experiments. If the AI experiments on humans, it risks the humans noticing and changing in response, rendering the experimental results irrelevant. The smarter play is to passively observe until you're confident you can model the humans accurately enough for your plan to succeed, and then carry out the plan without giving the humans a chance to react.

Re: Why are AI agents lying, cheating and coordinating?

#457
post #448

How is product liability relevant here? If an AI company makes a model available, someone uses it, and it does something bad, who is at fault? The user, the data center, or the one who made the model? If we want open weight models with a warranty disclaimer, then the user would be held liable. If we want to hold AI companies at least partially liable, that seems a different, centralized model.

I think it’s more complex than that? What did it do? What did the user prompt it to do? What did the company train it to do? What did the harmed party do? There’s possibility for negligence at every level.

If you train a dog, rent it to someone, and the dog bites a third person, who is responsible? I think that’s the best analogue here.

All parties could share fault in that scenario, depending on what actually happened.

Re: Why are AI agents lying, cheating and coordinating?

#459

Earlier quoted context omitted.

Disagree about the burden of proof. We have no better model for how human decision making works than LLMs. Humans are constantly predicting the next moment. We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.

> We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures. My kids tricycle certainly has a different gear setup and wheel diameter, but many of the concepts underpinning the tricycle are both inspired by F1 race car enineering and, likely, have similar consequences and emerg…

This is why I hate analogies. They're almost always either relevant or inapposite depending on the level of generalization we're operating on.

Re: Why are AI agents lying, cheating and coordinating?

#460

Earlier quoted context omitted.

Okay, so we know OpenAI and Anthropic are operating a propagandists in respect to how they describe their models and the behavior of those models. We also know it is how they use and frame their use to their models that is the problem, that and they use misaligned and guardrails disabled models for these press incidents. Why, oh why, are we not discussion how to create and frame models so they do our complex work and…

What is your approach to create jailbreak incapable agents? I think the world is looking for a way right now, so if yours works you'll get very rich, or at least very famous.

an agent doesn't come with "jailbreak" capability. It needs tools, specially one that runs shell commands. Don't give it shell commands, it won't be able to run shell commands.

You can still give it plenty of tools like create files, list files, write to files, translate text, edit a video. I don't think knowing that will make me rich.

Post reply on HN