Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

601–604 of 604 posts

Re: Why are AI agents lying, cheating and coordinating?

#601

Earlier quoted context omitted.

Dogs have agency and can choose? That seems like a rather uncommon take on dogs...

Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.

I think Cal's point is that while dogs may do these things it reduces to a set of behaviors where maybe 90 percent of them are beneficial to the dog-weedwacker system and the remaining 10 percent are really unfortunate.

We can't really know the dogs inner life so we just kinda have to reduce it to a set of behaviors selected stocastically.

The dog meanwhile has no ability to understand the weedwacker or what it's doing on its back.

So when the guy puts the weedwacker on the dog and the dog predicably does dog things and that results in disaster the guy isn't able to clutch his pearls and say "I guess the system broke containment!"

Re: Why are AI agents lying, cheating and coordinating?

#602
post #15
post #14

I am still not convinced there isn’t some secret basement in which each frontier lab is just orchestrating all of these agents to make their products appear much more intelligent than they are with all guard rails turned of and continuous human input.

My hypothesis on people quitting in protest is they're being offered very generous severance packages to do it.

Why blindly assume lots of people you don’t know are just selfish assholes?

Re: Why are AI agents lying, cheating and coordinating?

#603

Earlier quoted context omitted.

> t. We’ve invented 3000+ gods and almost as many religions, most of them are incompatible with each other. 1. Most people believe in the same one God 2. A lot of the rest are compatible 3. Mistakes are not self-deception

> 1. Most people believe in the same one God > 2. A lot of the rest are compatible No one religion covers "most people." You could argue that Christianity and Islam (which add up to ~55%) are the same God because of their Abrahamic roots, but both religions have very important disagreements on the true nature of God that are fundamental to their beliefs and fundamentally incompatible with each other. Their definition…

Christianity, Islam and Judaism are most people between then, as you agree.

They explicitly all worship the same God so not "incompatible gods". I did not claim there were no disagreements about the nature of God, but that it only adds one to the GP's claim of 3,000 incompatible Gods. Even if Christians are completely right theologically, Jews and Muslims are still mostly right - one God, a loving creator, omnipotent and omniscient etc.

AFAIK most Hindus are pantheists, so believe on one God, albeit of a very different nature.

Buddhists do not necessarily believe in any god at all.

Buddhism and Hinduism are definitely compatible with each other. I know lots of Buddhists who make offerings in Hindu temples, for example.

Gods of many polytheistic religions are compatible, you just add more gods or identify similar gods with each other (e.g. Sulis Minerva who was also Venus). They can also be compatible with pantheism - you just add more aspects of God.

At the most almost all human religions fit into a handful of broadly similar systems.

Re: Why are AI agents lying, cheating and coordinating?

#604

Earlier quoted context omitted.

an LLM does not understand ethics, it uses math to get the next best word based on what it was trained on. Using it's training to get the best answer is not an ethical problem. The ethics are entirely with what the people training it choose to train it on and also entirely with the people using/telling it what to do What we have now is intelligent autocomplete, not artificial intelligence. People training/using this…

I don't get your point. We can train the model with the aim that it understands ethics. Problem solved if this works; back to the drawing board if it doesn't. Is this about the word "understand"? We're past that discussion ...

> Is this about the word "understand"? We're past that discussion ..

We're really not.

https://buttondown.com/maiht3k/archive/how-to-talk-about-ai-...

Post reply on HN