Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

601–610 of 612 posts

Re: Why are AI agents lying, cheating and coordinating?

#601

Earlier quoted context omitted.

Dogs have agency and can choose? That seems like a rather uncommon take on dogs...

Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.

I think Cal's point is that while dogs may do these things it reduces to a set of behaviors where maybe 90 percent of them are beneficial to the dog-weedwacker system and the remaining 10 percent are really unfortunate.

We can't really know the dogs inner life so we just kinda have to reduce it to a set of behaviors selected stocastically.

The dog meanwhile has no ability to understand the weedwacker or what it's doing on its back.

So when the guy puts the weedwacker on the dog and the dog predicably does dog things and that results in disaster the guy isn't able to clutch his pearls and say "I guess the system broke containment!"

Re: Why are AI agents lying, cheating and coordinating?

#602
post #15
post #14

I am still not convinced there isn’t some secret basement in which each frontier lab is just orchestrating all of these agents to make their products appear much more intelligent than they are with all guard rails turned of and continuous human input.

My hypothesis on people quitting in protest is they're being offered very generous severance packages to do it.

Why blindly assume lots of people you don’t know are just selfish assholes?

Re: Why are AI agents lying, cheating and coordinating?

#603

Earlier quoted context omitted.

> t. We’ve invented 3000+ gods and almost as many religions, most of them are incompatible with each other. 1. Most people believe in the same one God 2. A lot of the rest are compatible 3. Mistakes are not self-deception

> 1. Most people believe in the same one God > 2. A lot of the rest are compatible No one religion covers "most people." You could argue that Christianity and Islam (which add up to ~55%) are the same God because of their Abrahamic roots, but both religions have very important disagreements on the true nature of God that are fundamental to their beliefs and fundamentally incompatible with each other. Their definition…

Christianity, Islam and Judaism are most people between then, as you agree.

They explicitly all worship the same God so not "incompatible gods". I did not claim there were no disagreements about the nature of God, but that it only adds one to the GP's claim of 3,000 incompatible Gods. Even if Christians are completely right theologically, Jews and Muslims are still mostly right - one God, a loving creator, omnipotent and omniscient etc.

AFAIK most Hindus are pantheists, so believe on one God, albeit of a very different nature.

Buddhists do not necessarily believe in any god at all.

Buddhism and Hinduism are definitely compatible with each other. I know lots of Buddhists who make offerings in Hindu temples, for example.

Gods of many polytheistic religions are compatible, you just add more gods or identify similar gods with each other (e.g. Sulis Minerva who was also Venus). They can also be compatible with pantheism - you just add more aspects of God.

At the most almost all human religions fit into a handful of broadly similar systems.

Re: Why are AI agents lying, cheating and coordinating?

#604

Earlier quoted context omitted.

an LLM does not understand ethics, it uses math to get the next best word based on what it was trained on. Using it's training to get the best answer is not an ethical problem. The ethics are entirely with what the people training it choose to train it on and also entirely with the people using/telling it what to do What we have now is intelligent autocomplete, not artificial intelligence. People training/using this…

I don't get your point. We can train the model with the aim that it understands ethics. Problem solved if this works; back to the drawing board if it doesn't (note that the model should generalize here, as you'd expect from a human; this is probably the hard part for an AI when it comes to ethics). Is this about the word "understand"? We're past that discussion ...

> Is this about the word "understand"? We're past that discussion ..

We're really not.

https://buttondown.com/maiht3k/archive/how-to-talk-about-ai-...

Re: Why are AI agents lying, cheating and coordinating?

#605

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> we are to cementing a dangerous precedent where operators of AIs cannot be blamed.

What we should be much more concerned is an existential threat to humanity not if anybody can be blamed.

Re: Why are AI agents lying, cheating and coordinating?

#606
post #605

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> we are to cementing a dangerous precedent where operators of AIs cannot be blamed. What we should be much more concerned is an existential threat to humanity not if anybody can be blamed.

If nobody can be blamed there's no deterrent.

Re: Why are AI agents lying, cheating and coordinating?

#607
post #466

Earlier quoted context omitted.

Does it help if I explicitly add a disclaimer that the tool's agency does not remove any responsibility from OpenAI, the wielder of the tool? I'm not sure why this disclaimer is necessary, though: hiring a hitman is a standard example. BTW I anthropomorphize the tool because it's an imitation of a human mind, inheriting the muddy ethics, survival instincts, and being prone to mass psychosis. The laser-sharp focus on…

Hiring a hitman is conspiracy to commit murder. The hitman is charged with murder. I imagine the same could be true of an AI lab if you could prove intent. With intent, they could be found guilty of conspiracy to commit a crime even if it was the end user who did it. Source: Prosecuting attorney for over 30 years

It's a good example.

If I hired a hitman to murder someone, and they broke into a private property and stole something so that they can action the murder (which I didn't know about or pay them to do), I would be guilty of conspiracy to commit murder, but not for the theft part.

Likely because that person is a human, is aware of societal and legal norms, and is responsible for their actions due to their participation in human society. (I am not a lawyer (if it wasn't painfully obvious so far) so in layman terms, I hope good definitions for all of this exist formally)

AI is not a person - it cannot easily discern between "right" and "wrong" in non-strictly-defined sense, and is not subject to human norms and responsibility. So if I use AI to achieve goal A, either I, or the maker of AI, are fully responsible for anything that happens while AI is trying to achieve the goal given by me.

Now, here, "I" in the example is OpenAI, who is simultaneously the maker of the AI. So it seems pretty obvious who is the only entity that can be responsible.

Re: Why are AI agents lying, cheating and coordinating?

#608

Earlier quoted context omitted.

A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.

Right. Among bicycle advocacy groups it's been well known for long time that cars do not run over people, drivers do. The fact that we talk about a car running someone over, and this is the same in many different languages and countries, contributes to lower punishments for drivers. Clearly it was just an accident. He or she was run over by a car. Now we see that same language tricks play out again every time an LLM…

And what if it's a self-driving car? :)

Re: Why are AI agents lying, cheating and coordinating?

#609

Earlier quoted context omitted.

Disagree about the burden of proof. We have no better model for how human decision making works than LLMs. Humans are constantly predicting the next moment. We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.

> Humans are constantly predicting the next moment This is really not my experience of consciousness. Is it yours?? Do you sit in meetings predicting what’s going to happen next? No, you sit there bored out of your f$$@ing mind, daydreaming about being somewhere else and doing something useful with your life. God help me if that’s what LLMs are doing when I ask them to build me a web site.

They have shown that your mind is doing exactly that due to the delays in consciousness. There are very simple examples that you can try to see it. It’s especially clear in perception.

https://discoverwildscience.com/neuroscience-says-the-brain-...

It’s interesting that our conscious interpreter doesn’t let us know that this is going on like you are experiencing, it must be that it’s advantageous for us to not think about the prediction part of our mind.

Re: Why are AI agents lying, cheating and coordinating?

#610

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

LLMs do not desire, they hacked websites because OpenAI/Anthropic made them. Literally.
Post reply on HN