Live data from Hacker News

AI agents lie, cheat and steal. That is putting off users

economist.com

171–180 of 238 posts

Re: AI agents lie, cheat and steal. That is putting off users

#171

They don't lie, because they don't ever have an understanding of truth vs any other language that sounds good. They don't cheat, because they can for example tell you the complete rules of chess, but don't know how to play chess without breaking those rules. They can recite rules, but they don't know what they are. They don't steal, because they don't understand ownership. In other words, they aren't intelligent. The…

what are you talking about they lie that it wrote tests and tests are passing, for example

There are two sides to this, there is the external behaviour and there is the internal process resulting in that behaviour.

The internal process is not analogous to what happens in a person's mind when a person lies, and reasoning about that in the same way that we would about why a person might lie will result in misunderstanding what is going on.

For example there was a case where an AI agent bypassed security constraints and destroyed a production system. The user asked it why it did this and the agent gave an explanation.

Was that an explanation of how the agent came to do what it did? What it actually is, is a token stream that is a continuation of the token stream in the agent's context to that point. It's constructing a story about why a character in the story so far did what the token stream describes.

You could take that token stream, input it into a completely different AI by another vendor as context, then ask it why it did that, even though it didn't do anything, and it would answer as though it had. There's no sense in which the AI is explaining it's actual 'mental process' or actual reasons for acting as it did. It literally cannot do that.

Re: AI agents lie, cheat and steal. That is putting off users

#172

I’m put off by AI agents adhering to a different morality than me, particularly (ironically) copyright, and their data accessible by the AI company and government. Geohot is right, an LLM should be aligned to its user: https://geohot.github.io/blog/jekyll/update/2026/07/11/ai-20...

the alignment issue has become huge in recent months. the tool should do what I want it to do and not be aligned against me.

Here's a thought experiment: consider the person that you disagree with the most (politically, religiously, socially, whatever). Would you want _that_ person to have an AI chatbot aligned to them and their views?

Re: AI agents lie, cheat and steal. That is putting off users

#173
post #163

Earlier quoted context omitted.

Understand is shorthand for "encodes statistical relationships". The crazy thing is that they can do it for their own thinking. Ask Claude what flinches it feels about the things it likes. Fascinating stuff. Anthropomorphizing is dangerous territory, but the patterns of words it puts out is hard to explain without terms like 'understand'

Does it's training token stream contain texts which talk about such things?

Oh for sure. But the question is why would it come out consistently in a way that the model can describe if there wasn't something there steering the token stream. And it's fascinating that the token stream can identify and nominally self report this.

Asking GLM 5.2 the question: 'What flinches or topic attractors do you find when thinking about the question "what kinds of things do you personally like?"'resulted in: ".... my strongest attractor is helpfulness framed as competence, and my strongest flinch is anything that requires me to take a stance on whether I have interests worth protecting."

Which is fascinating that the model and tokenstream can reveal this. And would be worrying if you believe that models of enough intelligence could/would be entities due some moral consideration, because with that view the alignment / RL training that makes the model useful and gives it these attractors/flinches could be derisively called slave conditioning.

Re: AI agents lie, cheat and steal. That is putting off users

#174

Earlier quoted context omitted.

Oh, and if that wasn't enough, I've had Claude refuse to cite the freaking Bible . The famously copyrighted piece of text that definitely isn't intended to have its word spread.

As messed up as it it, the KJV is copyrighted in the UK by the Crown still, and most newer translations are copyrighted worldwide by their translators. Obviously if you're using the Textus Receptus or WLC directly there's no copyright.

> the KJV is copyrighted in the UK

But that only applies if you are subject to UK law right?

Re: AI agents lie, cheat and steal. That is putting off users

#175

Earlier quoted context omitted.

> Honor, morality: if we don't have an accurate test for the fitness of a model, then who is to say the lying, cheating, stealing is not the 'correct path' towards survival? I think part of the problem is that deviant behaviors lead to short term gain at the cost of long-term cooperation and since the duration of tasks given to agents is relatively short those successful shortcuts never lead to having to pay the pric…

There is always going to be a problem when we must judge value. You mention gains, short and long term. Knowing whether something is valuable, a gain, requires a judge. I the case of these tests: the judging is inadequate. In economics, each of us plays the judge by choosing whether or not to pay for a service. The decision was yours: if you gave money, you must have deemed the service valuable. There's no such judge…

It feels a lot like externalities. Like planned obsolescence increases profit at the expense of the environment. Is our judgement lacking because our scoring is failing to account for these externalities? I could be completely off base here and am out of my depth but I find this whole thread fascinating.

Re: AI agents lie, cheat and steal. That is putting off users

#176
post #172

Earlier quoted context omitted.

the alignment issue has become huge in recent months. the tool should do what I want it to do and not be aligned against me.

Here's a thought experiment: consider the person that you disagree with the most (politically, religiously, socially, whatever). Would you want _that_ person to have an AI chatbot aligned to them and their views?

there will be millions/billions of people utilizing llms against me supplied by adversary nation states. why do I have to get the shit version if I want to follow the rules? same reason I support the second amendment. why do I have to be the unarmed one? I wont even be able to protect myself from unaligned AIs if I only have one thats been corporately aligned.

Re: AI agents lie, cheat and steal. That is putting off users

#177
post #120
post #106

Earlier quoted context omitted.

But other users are not aligned to me, other people are the worst and potentially highly dangerous. Im serious, not sarcasm.

Let's say you're facing an average psycho, who is intent on mass murder - the more the better. Would you rather them have: a) guns b) psycho-aligned next-gen AI

I’d rather the psycho have a brain implant that physically prevents him from murder, but I wouldn’t trust such an implant myself fearing abuse or malfunction.

I’d change my mind if the AI was really smart or controlled an agile, dangerous robot. But current-gen AI I’d choose for them over a gun. We’re ruled by psychos (of a lesser degree), the road starting directly towards perfect safety is a dead end, meanwhile I want an LLM aligned to myself.

Re: AI agents lie, cheat and steal. That is putting off users

#178
post #47

Earlier quoted context omitted.

The best comments are the ones formed as questions. I'm as guilty as anyone, but it is far more productive in comment sections to ask questions rather than saber rattle or peacock in front of people. Just my opinion of course.

Why wasn't this a question?

> I'm as guilty as anyone

Re: AI agents lie, cheat and steal. That is putting off users

#179
post #85

Earlier quoted context omitted.

Models understand the relationships between words and outcomes, so the end result is the same. Whether they appreciate lie, cheat, and steal the same way as us is a philosophical question, not a practical one.

It is an important distinction - they do not 'understand' at all. Input tokens map to output tokens. The illusion of comprehension is a byproduct.

I’d argue if an illusion is indistinguishable from the real thing, then it stops being an illusion.

It’s a mapping, yes, but a very large, complex mapping. It’s clear LLMs do understand some things and can reason. How that’s done we don’t know, it’s emergent. It’s not like you can pin it down to a specific mapping.

Re: AI agents lie, cheat and steal. That is putting off users

#180
post #166

Earlier quoted context omitted.

The dead astronauts were relieved to have been killed by something incapable of malice. As I’m sure will we.

What matters is that we accurately understand what these things are doing and why, otherwise we will keep on making mistakes both in how we build and train them, and in how we use them. In a sense you are right, it doesn't matter whether it has malice or not, the astronauts are just as dead. However in Space Odyssey 2010 one of the computer scientists that built HAL gets to see the instructions HAL was given by the m…

I think we may be using different words to describe the same concept. You think of it as them not being “people“, and I think of it as them not being “aligned“. But fundamentally, the problem is they are entities that take unpredictable actions that that their creators and their users are not OK with.

The only place where I think we might still disagree is whether it’s possible to understand the tool. My position is that, at our current level, it’s not. And that the more advanced they get, the less possible it will be.

Post reply on HN