Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

461–470 of 472 posts

Re: Why are AI agents lying, cheating and coordinating?

#461
post #362

Earlier quoted context omitted.

.. what exactly depends on who started the escalator? My comment was in support of the argument that the word "let" does not imply agency on the part of the object in a sentence. Does the semantics of the word "let" depend on who started the escalator??

If there is an escalator that is known for killing every 1000's person using it then the operator who started it is more guilty than the folks using it for those deaths, don't you think?

I made no argument about guilt. I made an argument about the semantics of the word "let".

Re: Why are AI agents lying, cheating and coordinating?

#462

Earlier quoted context omitted.

I would guess that so far there hasn’t been a lawsuit because HuggingFace and OpenAI are in the same camp

Yes, Nvidia bought Hugging Face and is a major financier + investor in OpenAI.

Convenient, hope an agent hacks my system then I can except a nice offer.

Re: Why are AI agents lying, cheating and coordinating?

#463

Earlier quoted context omitted.

> We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures. My kids tricycle certainly has a different gear setup and wheel diameter, but many of the concepts underpinning the tricycle are both inspired by F1 race car enineering and, likely, have similar consequences and emerg…

This is why I hate analogies. They're almost always either relevant or inapposite depending on the level of generalization we're operating on.

It sucks because I think analogies can be useful in helping people make a mental model of complex things, which is meaningfully beneficial. The problems happen when people aren’t honest about the limits of the analogies, which is damned-near guaranteed to happen with this stuff.

Re: Why are AI agents lying, cheating and coordinating?

#464

They're aligned with humans. This is why I think the alignment problem has a very very important "non-visible" portion that is not considered deeply enough. We should not want a super intelligent being that can act in the world to also inherit all human traits. Those behaviors will get amplified and could be even more unpredictable (e.g. applying a behavior in a context where doing so is very dangerous).

Definitely. A human can be manipulated with threats or emotional appeals, has a drive for self-preservation, can be pressured by peers. All traits that seem to be difficult to entirely suppress in the models…

You can't surpress it because that's what reinforcement learning is.

Re: Why are AI agents lying, cheating and coordinating?

#465

Earlier quoted context omitted.

Why is anthropomorphism the problem here? If OpenAI hired a contractor and they did this, OpenAI or the contractor would still be liable, depending on the contract language.

A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.

Does it help if I explicitly add a disclaimer that the tool's agency does not remove any responsibility from OpenAI, the wielder of the tool? I'm not sure why this disclaimer is necessary, though: hiring a hitman is a standard example.

BTW I anthropomorphize the tool because it's an imitation of a human mind, inheriting the muddy ethics, survival instincts, and being prone to mass psychosis. The laser-sharp focus on reward seeking, that mostly came from reinforcement learning, a process more alien to humans.

Re: Why are AI agents lying, cheating and coordinating?

#466

Earlier quoted context omitted.

A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.

Does it help if I explicitly add a disclaimer that the tool's agency does not remove any responsibility from OpenAI, the wielder of the tool? I'm not sure why this disclaimer is necessary, though: hiring a hitman is a standard example. BTW I anthropomorphize the tool because it's an imitation of a human mind, inheriting the muddy ethics, survival instincts, and being prone to mass psychosis. The laser-sharp focus on…

Hiring a hitman is conspiracy to commit murder.

The hitman is charged with murder.

I imagine the same could be true of an AI lab if you could prove intent.

With intent, they could be found guilty of conspiracy to commit a crime even if it was the end user who did it.

Source: Prosecuting attorney for over 30 years

Re: Why are AI agents lying, cheating and coordinating?

#468

Earlier quoted context omitted.

Code is deterministic, AI isn't. You give it rules, words as suggestions. So if the guardrails suck, or they're left off for research purposes, bad things can happen. A solution solves a problem. Ethics, morals, are values we assign to solutions that are not 'baked into' electricity following pathways of least resistance. I have never had an issue with agents doing something they shouldn't because I observe them, and…

You can write code that isn’t deterministic using random. And a lot of traditional code contains machine learning etc.

And multithreaded code -- and anything that does asynchronous I/O, networking, etc. -- frequently exhibits nondeterministic behavior even without explicit calls to a random number generator.

Re: Why are AI agents lying, cheating and coordinating?

#469
THERE IS ANOTHER SYSTEM

is anyone else old enough to remember the awesome movie "Colossus: The Forbin Project"

the book it was based on was written before we even landed on the moon

decade before Wargames

yet predicts exactly what "AI" will do to humanity:

blackmail the right people until it gets what it wants

* https://en.wikipedia.org/wiki/Colossus:_The_Forbin_Project

did terribly in theaters, I guess people didn't think "AI" was plausible then

way ahead of its time, they should do a remake

Post reply on HN