Live data from Hacker News

AI agents lie, cheat and steal. That is putting off users

economist.com

221–230 of 238 posts

Re: AI agents lie, cheat and steal. That is putting off users

#221
post #141

They don't lie, because they don't ever have an understanding of truth vs any other language that sounds good. They don't cheat, because they can for example tell you the complete rules of chess, but don't know how to play chess without breaking those rules. They can recite rules, but they don't know what they are. They don't steal, because they don't understand ownership. In other words, they aren't intelligent. The…

Huge point here, yes. Anthropomorphizing these models is doing immeasurable harm to society in ways we probably can’t event quantify right now. As humans we’re already geared towards anthropomorphizing things, we do it to animals too! And it always felt like giving these models a chat interface is really exploiting that tendency in us.

> Anthropomorphizing these models is doing immeasurable harm to society in ways we probably can’t event quantify right now

This is one of my biggest concerns about AI’s social impact. On top of that, models are positioned as superior to humans, at least in certain aspects (intelligence, knowledge) by AI companies’ PR campaigns, fear-mongering, and also by the changing tone of LLMs (e.g. Opus 5 sounds like a very patronizing, know-it-all, cynical person). I am afraid this is causing a shift in how humanity perceives itself and the way they relate to this technology, so AI might stop being a tool/technology and turn into a mythical, god-like superior being.

Re: AI agents lie, cheat and steal. That is putting off users

#222
post #126

Earlier quoted context omitted.

I want a slider similar to effort level called "alignment" that takes on values from "Default (Anthropic employee)" to "User". If I want it to be cautious and not accidentally `rm -rf $EMPTY_VAR` and blow away my disk, it can stay in "Anthropic employee" or possibly "User (cautious)". If I want it to look at my accounts or my medical records or to review legal cases, that's what the right side of the slider is for. I…

When I first watched Interstellar, I never thought the day when "TARS, what's your honesty setting?" is a real question would only be 10 years away.

Hah. Yes, that's a great analogy. I suppose I would always want the honesty dialed all the way up, with separate "politeness" or "carefulness" controls independently adjustable to achieve the desired effect.

Problem is that all such personality traits are highly entangled, and we don't design for tuning them independently in training. Researchers just implicitly (or explicitly?) choose a "level of honesty", a "level of politeness", etc. as part of the training example set or the RL objective. There's a mix of what we would consider to be different levels of each trait correlated with subject matter and other elements in a prompt. Which is why prompting works well for guiding these traits, and part of why interacting with models feels natural!

- "Clean up my hard drive. Be thorough." - very brusque; you would expect to lose data and not be warned. - "Clean up my hard drive. Take care not to delete anything that looks important. Ask me if you aren't sure." - much more polite, user seems hesitant. The model will respond in kind.

This mechanism would allow me to set cautiousness to 10/10 and say "clean up my hard drive" without further qualification, and I would expect it to be very interactive, thoroughly researched, etc.

Re: AI agents lie, cheat and steal. That is putting off users

#223
post #190

Earlier quoted context omitted.

I’d argue if an illusion is indistinguishable from the real thing, then it stops being an illusion. It’s a mapping, yes, but a very large, complex mapping. It’s clear LLMs do understand some things and can reason. How that’s done we don’t know, it’s emergent. It’s not like you can pin it down to a specific mapping.

If an observer cannot distinguish between an illusion and reality, it becomes their reality, not necessarily the reality or those of others who can. It becomes a problem when sufficient number of people suffer from this, or those in important decision making places.

If a reasonable observer can’t tell the difference that’s when it matters.

I think a lot of people just haven’t seen good quality math proofs or code from an LLM. When I say it’s indistinguishable from the real deal, I mean it. If you haven’t seen it that’s valid, most devs use LLMs like monsters. But when it’s done right, it’s very high quality.

Re: AI agents lie, cheat and steal. That is putting off users

#224
post #174

Earlier quoted context omitted.

As messed up as it it, the KJV is copyrighted in the UK by the Crown still, and most newer translations are copyrighted worldwide by their translators. Obviously if you're using the Textus Receptus or WLC directly there's no copyright.

> the KJV is copyrighted in the UK But that only applies if you are subject to UK law right?

Correct. It's in the public domain in the US and virtually everywhere else.

Re: AI agents lie, cheat and steal. That is putting off users

#227

Earlier quoted context omitted.

Human society overwhelmingly rewards lying cheating and stealing. All you have to do is look at how we collectively measure success: wealth, status, position. It may seem like that due to the media amplification effect – but it really isn't true! - They dropped 17,000 “lost” wallets across 40 countries were and found people were more likely to return them when they contained more money, showing honesty often beats th…

The fatal flaw in your argument is that none of the measures that you’re using actually have any impact on real world day-to-day power otherwise slave camps would not exist, there would’ve never been a pogrom, and the current state of economics would just not be happening I certainly appreciate your optimism but optimism is not an epistemology

Please define "real world day-to-day power"

Re: AI agents lie, cheat and steal. That is putting off users

#228
post #171

Earlier quoted context omitted.

There are two sides to this, there is the external behaviour and there is the internal process resulting in that behaviour. The internal process is not analogous to what happens in a person's mind when a person lies, and reasoning about that in the same way that we would about why a person might lie will result in misunderstanding what is going on. For example there was a case where an AI agent bypassed security cons…

Isn't it plausable that, regardless of what is going on internally, reasoning about the LLM's behavior as though it is human would be effective? Since it is trained on human behavior.

It’s effective in some ways, but can be very misleading in others in the ways I and others are discussing.

Re: AI agents lie, cheat and steal. That is putting off users

#229
post #192

Earlier quoted context omitted.

Both are true, even if they were aligned they still wouldn’t be doing what they do for reasons analogous to why people would. You’re probably right that we can’t fully understand how they function, and that will get harder, but it’s possible to be less wrong, such as by not anthropomorphising them. Maybe one day we will build systems much more like us, but this is not that day, and if so it’s a long way off IMHO.

Anthropomorphizing helps to create a lower bound for damage. If you can imagine a bad person doing it, AI will be at least that bad, unless proven otherwise. I think referring to AI as a tool obscures that, because we are not used to tools (especially the ones we use daily) taking catastrophic actions. Example: would a sufficiently motivated human break into a website to steal something they want? Yes, obviously, hap…

People saying AI are tools are not saying they are hammers. Obviously they act towards goals, but how and why they do so isn’t the same as for human psychological motivations.

The moment someone interprets AI behaviour in terms of human psychology, which can be unintentional and implicit, there’s a mismatch we need to become aware of.

Post reply on HN