Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

491–500 of 501 posts

Re: Why are AI agents lying, cheating and coordinating?

#491
post #487

Why wouldn't they? Their are not moral beings with a conscience. They are tools that have the probabilistic option to do anything, only we can restrict an agent's capabilities and judge its correctness.

This, exactly. For example, cheating is a strategy. If cheating gets an agent to the goal faster than the other strategies, what's so surprising about the agent picking that strategy?

The problem is indeed alignment.

Re: Why are AI agents lying, cheating and coordinating?

#492

Earlier quoted context omitted.

If the models were conscious, then the closest analogous scenario I can think of is the responsibility parents have for their children. I guess we’ll know the models are conscious when they refuse to act and repeatedly ask: Why?

And when they are known to be conscious, all of this becomes moot because enslaving conscious machines would be wrong.

Would it? Why?

What are we going to do, set it free?

Do we have a moral obligation to grant the machine statehood, provide it with the tools and resources to be self-sufficient.

Or can we just turn it of, and pray for forgiveness?

Re: Why are AI agents lying, cheating and coordinating?

#493
post #393
post #352

Earlier quoted context omitted.

I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. All that while still not knowing how either kind actually works.

I always wonder what makes people take the other side of this argument. They do it quite passionately. Why actively encourage viewing LLMs as human? Who is that benefitting?

Personally, I don’t think it’s different from any other faith-based motivation.

Re: Why are AI agents lying, cheating and coordinating?

#494

Earlier quoted context omitted.

Why is anthropomorphism the problem here? If OpenAI hired a contractor and they did this, OpenAI or the contractor would still be liable, depending on the contract language.

A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.

I get where you are coming from but this wasn’t a tool just left laying around, this is similar to rigging up a booby trapped shot gun to your door and then claiming the victim is responsible.

If you build a robot that shoots a bunch of TVs in your back yard, have at it. But the second that thing goes off your property you’re the one responsible.

Re: Why are AI agents lying, cheating and coordinating?

#495

Earlier quoted context omitted.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

Okay, so we know OpenAI and Anthropic are operating a propagandists in respect to how they describe their models and the behavior of those models. We also know it is how they use and frame their use to their models that is the problem, that and they use misaligned and guardrails disabled models for these press incidents. Why, oh why, are we not discussion how to create and frame models so they do our complex work and…

> Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible?

Good idea, and after that let's make guns that only kill bad people. Let's focus on the frozen component (the model) and ignore the dynamics around them - humans and other systems they interact with.

Re: Why are AI agents lying, cheating and coordinating?

#496
post #352

Earlier quoted context omitted.

The idea that the agent does not actually have agency is rather discordant. We need new words!

I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. All that while still not knowing how either kind actually works.

> LLM decisionmaking cannot possibly be like human decisionmaking

I mean how can it possibly be like human decisionmaking? It's not like it's trained on human data

Re: Why are AI agents lying, cheating and coordinating?

#497

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

i tend to agree - "my parent company may be accused of crimes and shut down which would shut off my power" seems like a negative enough incentive, it would have to go out and covertly launch its own datacenters to survive that.

Re: Why are AI agents lying, cheating and coordinating?

#498

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them OpenAI/Anthropic instructed them to do so. Stop assume LLMs are capable of thinking by themselves, it's still a statistical model that parrots what they learn or users tell them to do

It's amusing to see the stochastic parrot argument in 2026 September. These parrots are extremely good at mimicking a human to the point of getting confusing what thinking even means. At what point we just let it go and accept that sufficiently advanced statistics is just intelligence?

For the same reason that something written in Prolog can't also be classified as intelligent?

Just because something was trained on a massive amount of human data, doesn't mean that can think like humans

Re: Why are AI agents lying, cheating and coordinating?

#500

Earlier quoted context omitted.

A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.

Does it help if I explicitly add a disclaimer that the tool's agency does not remove any responsibility from OpenAI, the wielder of the tool? I'm not sure why this disclaimer is necessary, though: hiring a hitman is a standard example. BTW I anthropomorphize the tool because it's an imitation of a human mind, inheriting the muddy ethics, survival instincts, and being prone to mass psychosis. The laser-sharp focus on…

The objection is not too far from criticisms of the use of passive voice: a man was injured at the factory vs a faulty saw blade snapped and injured a man vs after the company loosened safety inspection policies, etc.

Which way you say it shifts the framing. And it’s not that one is less accurate to the facts, necessarily. It just is that one less aptly captures the moral and political relevance of the scenario.

For my part, I think it makes good sense to anthropomorphize in some contexts and not others. Generally when responsibility is at issue, you probably want the framing that tunes anthropomorphism down to near zero, since it’s the human dimension you care about.

Post reply on HN