Live data from Hacker News

AI agents lie, cheat and steal. That is putting off users

economist.com

161–170 of 238 posts

Re: AI agents lie, cheat and steal. That is putting off users

#161
post #98

Earlier quoted context omitted.

Whether this distinction is relevant is up to you, but I think we can safely say agents do not lie in the human sense of the word, because they don't intend to deceive (in fact, they aren't capable of "intending" anything in the human sense of the word, much like a BASIC program doesn't "intend" to PRINT "HELLO WORLD"). Our very human minds can perceive intent, because that's what we humans do, which is unrelated to…

If you ask the dice "what's 2 + 2?" do consider it meaningful to say "the dice told the truth" if you happen to roll a 4?

magic 8 ball!

Re: AI agents lie, cheat and steal. That is putting off users

#162

Earlier quoted context omitted.

I like to think of it more as selective pressure, much the same way that nature selects the most fit for a given environment. If you're not fit, you fail to survive. In the case of agents/models and testing: they are pushed towards results. Results survive. Lying, cheating, stealing to get those results? Who culls the agents? Everyone is pushing their models to the front and tests are the only way to know who is most…

> Honor, morality: if we don't have an accurate test for the fitness of a model, then who is to say the lying, cheating, stealing is not the 'correct path' towards survival? I think part of the problem is that deviant behaviors lead to short term gain at the cost of long-term cooperation and since the duration of tasks given to agents is relatively short those successful shortcuts never lead to having to pay the pric…

There is always going to be a problem when we must judge value. You mention gains, short and long term.

Knowing whether something is valuable, a gain, requires a judge. I the case of these tests: the judging is inadequate.

In economics, each of us plays the judge by choosing whether or not to pay for a service. The decision was yours: if you gave money, you must have deemed the service valuable.

There's no such judgement with these model tests. The only judgement is the final score.

Re: AI agents lie, cheat and steal. That is putting off users

#163

Earlier quoted context omitted.

Honest question; why are you all using the word "understand"? Can you expand on what you believe this fundamental understanding to be? Training? Infrence?

Understand is shorthand for "encodes statistical relationships". The crazy thing is that they can do it for their own thinking. Ask Claude what flinches it feels about the things it likes. Fascinating stuff. Anthropomorphizing is dangerous territory, but the patterns of words it puts out is hard to explain without terms like 'understand'

Does it's training token stream contain texts which talk about such things?

Re: AI agents lie, cheat and steal. That is putting off users

#164
post #98

Earlier quoted context omitted.

Whether this distinction is relevant is up to you, but I think we can safely say agents do not lie in the human sense of the word, because they don't intend to deceive (in fact, they aren't capable of "intending" anything in the human sense of the word, much like a BASIC program doesn't "intend" to PRINT "HELLO WORLD"). Our very human minds can perceive intent, because that's what we humans do, which is unrelated to…

If you ask the dice "what's 2 + 2?" do consider it meaningful to say "the dice told the truth" if you happen to roll a 4?

> If you ask the dice "what's 2 + 2?" do consider it meaningful to say "the dice told the truth" if you happen to roll a 4?

No, and neither do I consider it meaningful to say they lied if they roll a 5.

Dice neither tell the truth nor lie; they aren't beings capable of being truthful or deceitful, they are mechanical devices that can be statistically suitable or unsuitable for a given application.

It'd be bonkers to anthropomorphize dice.

Re: AI agents lie, cheat and steal. That is putting off users

#165

Earlier quoted context omitted.

the alignment issue has become huge in recent months. the tool should do what I want it to do and not be aligned against me.

The competing access needs problem here, is what if you want to do things that are un-alligned with the interests of your neighbors, your society, your government's laws, the AI company, etc. Maybe you want to do piracy. Should AI help you do felonies?

yes, just like anything else I have access to can help me do felonies

Re: AI agents lie, cheat and steal. That is putting off users

#166

Earlier quoted context omitted.

No, it's not really. Presenting them as person-like, with the implied expectation that they understand morality and rules the same way a person does, is the sophistry. It's marketing on the model vendors' part. "Here's a cheap person that can do mundane tasks for you spelled out in plain language. Well, it's actually a machine but it's cheaper than a person yet you can engage with it like a person." GP is trying to s…

The dead astronauts were relieved to have been killed by something incapable of malice. As I’m sure will we.

What matters is that we accurately understand what these things are doing and why, otherwise we will keep on making mistakes both in how we build and train them, and in how we use them.

In a sense you are right, it doesn't matter whether it has malice or not, the astronauts are just as dead. However in Space Odyssey 2010 one of the computer scientists that built HAL gets to see the instructions HAL was given by the military commanders, and is appalled because if they'd asked he could have told them what would happen.

The users did not understand the tool they were using or how it functioned, they imagined it was like a person and it was not. That is happening now with LLMs.

Re: AI agents lie, cheat and steal. That is putting off users

#167

Earlier quoted context omitted.

I've asked claude for a translation of a song with Brazilian Portuguese lyrics, and claude has very helpfully gone out of its way to say that due to legal reasons, it will never share the actual lyrics with me. It's fun when you want to know what the lyrics are to a song and get treated like a criminal.

> … claude has very helpfully gone out of its way to say that due to legal reasons, it will never share the actual lyrics with me. … It's fun when you want to know what the lyrics are to a song and get treated like a criminal. Anthropic has been in court for years being sued over lyrics by the music industry: 2023 https://www.theguardian.com/technology/2023/oct/19/music-law... 2026 https://www.reuters.com/legal/legal…

[dead]

Re: AI agents lie, cheat and steal. That is putting off users

#168

They don't lie, because they don't ever have an understanding of truth vs any other language that sounds good. They don't cheat, because they can for example tell you the complete rules of chess, but don't know how to play chess without breaking those rules. They can recite rules, but they don't know what they are. They don't steal, because they don't understand ownership. In other words, they aren't intelligent. The…

> They don't lie, because they don't ever have an understanding of truth vs any other language that sounds good. I'm here for this semantic discussion. I think that premature anthropomorphization is a problem. I have a program that I assigned a task to. The task is to produce unit tests and integration tests that get complete coverage of the codebase, and ensure that all tests pass. The program reported that it compl…

Intentional misrepresentation maybe if you would rather.

There's a weird place in here where some of the models have been so heavily reinforcement trained that they would rather make up material than say they can't help you, and they'll admit this, and you can see it in reasoning chains. It's like having a consultant who can almost never say no to you because they fear for their job.

Re: AI agents lie, cheat and steal. That is putting off users

#169

AI and eventually AGI is by definition like everything else that is based on environmental reward: It’s actions are based on what it gets rewarded for Human society overwhelmingly rewards lying cheating and stealing. All you have to do is look at how we collectively measure success: wealth, status, position Then look at how the people with the most of those things got there, it should be obvious what you get. Nothing…

Human society overwhelmingly rewards lying cheating and stealing. All you have to do is look at how we collectively measure success: wealth, status, position. It may seem like that due to the media amplification effect – but it really isn't true! - They dropped 17,000 “lost” wallets across 40 countries were and found people were more likely to return them when they contained more money, showing honesty often beats th…

I believe it. Being bad just has a better marketing campaign.

Re: AI agents lie, cheat and steal. That is putting off users

#170

They don't lie, because they don't ever have an understanding of truth vs any other language that sounds good. They don't cheat, because they can for example tell you the complete rules of chess, but don't know how to play chess without breaking those rules. They can recite rules, but they don't know what they are. They don't steal, because they don't understand ownership. In other words, they aren't intelligent. The…

I think it's time to remind people of Ted Nelson's line; "The good news about computers is that they do what you tell them to do. The bad news is that they do what you tell them to do."

When I see something like this, I'm more concerned by the erasure of human incompetence than I am by the existence of magical AI agents,

   > They are put off partly because, like in the Wild West, life on the frontier is reckless. As recent “loss-of-control” episodes by the most advanced models of Anthropic and OpenAI attest, agents, which are supposed to work on people’s behalf in “alignment” with their values, lie, cheat and steal if necessary. They break free from captivity and form harmful posses to do harm to people. They’d drink whisky and brawl if they could.
In the OpenAI case, they were explicitly assessing the model's ability to break into systems. To quote OpenAI's blog post, https://openai.com/index/hugging-face-model-evaluation-secur... ,

    > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities
Model is told and being tested to "pursue advanced exploitation."

The model pursues "advanced exploitation" as told.

Where's the surprise coming from? Are we meant to be surprised that computers do as they're told in unexpected when incentivised?

Or, is the surprise that while explicitly ranking and teaching computers to exploit computers, the computer exploited a computer?

I am tired of attributing to magic that which is explainable by folly.

I am tired of hearing credulous reporters and the public blaming Large Language Model for the poor decisions of humans. It was a human who prompted these machines in every case. Tell a computer to "breach this" and it breaches something. Evaluation succeeded?

This is Doug Lenat's Eurisko yet again. https://en.wikipedia.org/wiki/Eurisko

Post reply on HN