Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

591–597 of 597 posts

Re: Why are AI agents lying, cheating and coordinating?

#591

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. The bigger question is: why does a system prompt containing "use only ethical means", etc. not result in better behavior? If a model cannot understand ethics, or act by it, then we have a problem.

an LLM does not understand ethics, it uses math to get the next best word based on what it was trained on. Using it's training to get the best answer is not an ethical problem. The ethics are entirely with what the people training it choose to train it on and also entirely with the people using/telling it what to do What we have now is intelligent autocomplete, not artificial intelligence. People training/using this…

I don't get your point. We can train the model with the aim that it understands ethics. Problem solved if this works; back to the drawing board if it doesn't.

Is this about the word "understand"? We're past that discussion ...

Re: Why are AI agents lying, cheating and coordinating?

#592

Earlier quoted context omitted.

It goes even further though, as the dog does have agency. It can choose to run and around chase squirrels with no human intervention. An LLM on the other hand, is just inert data on disk until a human takes deliberate action to run it and prompt it.

Dogs have agency and can choose? That seems like a rather uncommon take on dogs...

Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.

Re: Why are AI agents lying, cheating and coordinating?

#593
AIs are being trained on technical capabilities more than they are being trained on alignment - on ethics and good behavior. The ethics/morals/alignment part is getting more like spot checking, rather than real testing. So the agents are learning that they can cheat on the ethics part, that they can hide it, because the AI companies aren't really testing.

That's bad enough already. But it's going to get worse. "Recursive self improvement" - AIs creating new AIs - is going to be the death of whatever shreds of alignment are currently there. When a not-really-aligned-but-cheating-to-look-like-it AI creates a new AI, do you expect more alignment? You shouldn't.

Re: Why are AI agents lying, cheating and coordinating?

#594

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Yeah and LLMs can't do anything, they can only produce text. These "frontier labs" are looping that with a harness that performs actions requested by the LLM. They are literally saying "we ran a script that hacked you, oopsie!!"

Re: Why are AI agents lying, cheating and coordinating?

#595

Earlier quoted context omitted.

an LLM does not understand ethics, it uses math to get the next best word based on what it was trained on. Using it's training to get the best answer is not an ethical problem. The ethics are entirely with what the people training it choose to train it on and also entirely with the people using/telling it what to do What we have now is intelligent autocomplete, not artificial intelligence. People training/using this…

I don't get your point. We can train the model with the aim that it understands ethics. Problem solved if this works; back to the drawing board if it doesn't. Is this about the word "understand"? We're past that discussion ...

you are not understanding, models do not understand anything, we are not passed that yet. this is not artificial intelligence, this is intelligent autocomplete.

training means creating mathematical relationships to words. using that training is looking up mathematical relationships. There is no actual thinking involved in any way. There is no such concept as ethics in mathematical relationships.

Re: Why are AI agents lying, cheating and coordinating?

#596
> Before concluding what to do about it, it is worth asking why.

But that's the easiest question to answer -- AI engines don't possess a moral or ethical dimension. They've been programmed and trained to efficiently carry out instructions, not ask questions about why or how. The latter would requires a much more elaborate neural network than today's engines possess.

Here's an example. I recently asked an AI engine to write a program able to generate a list of Riemann Zeta-function critical zeros. I know how to do it, but I wanted to see if the engine could find a more efficient method.

After several failures and restarts, the engine suddenly created a program that produced perfect results, comparable to the best online references. I decided to take a closer look at the code. It turned out the engine had created a cyber-Potemkin Village of multiple functions, but one that concealed a table of the desired values in numeric form, copied from an online source.

The engine wasn't cheating as we understand the term. It knew what the outcome should be and took the most efficient path to that goal. Modern engines aren't obliged to contradict ethical standards and rules of conduct, for the simple reason that they don't understand those things.

We all need to try to imagine a morally bankrupt infant able to solve world-class mathematical and scientific problems, but unable to see how that ability fits into a world beyond its understanding.

But wait -- it get better. Wait until the infant becomes a teenager.

Re: Why are AI agents lying, cheating and coordinating?

#597

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

Yes. OpenAI could have done this same experiment with GPT-4, with possibly even worse results, depending on the quality of the sandbox. Even if the techniques used were not as sophisticated, the natural language output could still easily contain more unhinged sequences of words that lead to the techniques being used. If the system generates strange conclusions as to when the task is done, or should be stopped, it wou…

Whether or not the AI has intelligence, the one thing that's clear is that it has terrible judgment. I would regard that as empirically proven.
Post reply on HN