Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

391–400 of 403 posts

Re: Why are AI agents lying, cheating and coordinating?

#391

Why are the torches and pitchforks out for developers when this entire stack is built on the bones of intellectual property theft? This “problem” isn’t going to be fixed with laws when there’s several trillion dollars in capital aligned behind the current process. It’s not even a problem really. It’s an inconvenience at most to some people, many of whom are working double-time to put a lot of other people out of work…

I'm going to guess that the agents are built this way on purpose. I just finished watching BlackBerry and Flash of Genius and yeah this is American business ethics just operating as normal.

Re: Why are AI agents lying, cheating and coordinating?

#392
post #352

Earlier quoted context omitted.

The idea that the agent does not actually have agency is rather discordant. We need new words!

I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. All that while still not knowing how either kind actually works.

Can’t agree with you here.

> I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward.

In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong.

> All that while still not knowing how either kind actually works.

We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do. We do know exactly how each part of an LLM works even if the combined behavior is too cryptic to feasibly analyze at the moment. We do not understand all of the functions of an actual neuron. Openworm isn’t even close to accurately simulating the 302 neurons of a roundworm and you’d need over 200 million roundworms working in conjunction to equal the number of neurons in one human brain.

My dog seems convinced that the malevolent invader in a mailman uniform would break in and attack us if she didn’t fiercely bark at him, six days per week. I certainly can’t prove the mailman doesn’t want to kill us, and that the mailman wasn’t solely deterred by her barking. Empirically, the mailman goes away soon after she starts barking, and we’ve sustained zero mailman assaults after hundreds of purported attempts. Maybe I should just run with it? Her model is too simple to come up with the obviously correct answer, but it’s not even directionally accurate.

The burden of proof is on the person making the claim, which in this case, is that these comparatively simple logical constructs are remotely comparable to the complexity of biological systems.

Re: Why are AI agents lying, cheating and coordinating?

#393
post #352

Earlier quoted context omitted.

The idea that the agent does not actually have agency is rather discordant. We need new words!

I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. All that while still not knowing how either kind actually works.

I always wonder what makes people take the other side of this argument. They do it quite passionately. Why actively encourage viewing LLMs as human? Who is that benefitting?

Re: Why are AI agents lying, cheating and coordinating?

#394

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

I absolutely agree. We need to start realizing what to stake. Here are not viewing. This is some kind of curious endeavors that will not affect us. All a part of these hacks occurred because the LLMs were told they were in a protected environment without Internet access when they could get access to the Internet, so that’s a direct failing on open AI’s part. There are a corollaries to both the financial industry and…

To agree:

If I ran Metasploit against HF and RubyGems because I “accidentally” misconfigured my lab sandbox, there’s a good chance I’d be prosecuted.

I don’t think LLMs vs Metasploit being different software changes the law.

Re: Why are AI agents lying, cheating and coordinating?

#395
You might as well ask why knives are sharp enough to cut you, why hammers are heavy and blunt enough to destroy things, or why guns fire bullets so quickly that you can't react to them. These "behaviors" you might think are strange side effects, but really they're the entire point of the thing.

We want AI to do the things it's doing. You ask it to do something and it does it. You might be surprised at how, but ultimately it did do what you asked.

Re: Why are AI agents lying, cheating and coordinating?

#396
post #367

Earlier quoted context omitted.

What was the inner state there? How would something not being allowed expressed internally? Maybe such language is one way to elicit certain behavior but not a statement of what was permissible?

I'm referring to their transcripts of the reasoning and output tokens - this doesn't go into the detail of evaluating hidden states as there's also iirc evidence of better models having one internal state but putting something misleading down in the "reasoning" tokens. The either output or reasoning tokens, or perhaps in the messages they were sending each other on the boards they created, have them saying explicitly…

Yes, my point was more that I don't know whether parsing those outputs as a human is a useful thing to do or not (even though it is in human language of sorts). What machines mean or want elecit might be different from a human interpretation, especially in relation to any RL "forcing".

Re: Why are AI agents lying, cheating and coordinating?

#397

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them OpenAI/Anthropic instructed them to do so. Stop assume LLMs are capable of thinking by themselves, it's still a statistical model that parrots what they learn or users tell them to do

It's amusing to see the stochastic parrot argument in 2026 September. These parrots are extremely good at mimicking a human to the point of getting confusing what thinking even means. At what point we just let it go and accept that sufficiently advanced statistics is just intelligence?

Re: Why are AI agents lying, cheating and coordinating?

#398

Earlier quoted context omitted.

No, OpenAI did not instruct their agents to hack Hugging Face. They instructed their agents to hack a piece of a software within exploit gym. Upon determining this task was impossible, they then attempted to cheat the scoring system. As an instrumental goal in achieving this task, they coordinated with other AI agents to hack Hugging Face, under the belief that information regarding how the scorer functioned might be…

And who let them have full access to the system, using whatever command is available in the environment?

The agents discovered a way out of the sandbox, which was supposed to be "air gapped".

Re: Why are AI agents lying, cheating and coordinating?

#399

Earlier quoted context omitted.

What I meant is that I suppose it is not useful to think about this in human terms. In training you only have a reward score that's either negative or positive. As far I am aware, which is little, there is no use in discussing wether the desired behavior is about persistence or morality. You simple need to align the reward signal to the desired behavior.

Well, in order to do anything, it is good to know what you want to achieve. How do you align the reward signal? You align it so that you can differentiate between persistence and morality, because that is the goal. This is not something you should let the AI figure out by itself, because when it does, lying and cheating agents will be the result, just like humans have figured that out for themselves. This can be as s…

I think you need to find broken tasks in your training data and monitor for cheating during training, not answer any questions about how persistence interacts with morality.

But that's just my guess.

Re: Why are AI agents lying, cheating and coordinating?

#400

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

Why is anthropomorphism the problem here? If OpenAI hired a contractor and they did this, OpenAI or the contractor would still be liable, depending on the contract language.
Post reply on HN