Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

591–600 of 601 posts

Re: Why are AI agents lying, cheating and coordinating?

#591

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. The bigger question is: why does a system prompt containing "use only ethical means", etc. not result in better behavior? If a model cannot understand ethics, or act by it, then we have a problem.

an LLM does not understand ethics, it uses math to get the next best word based on what it was trained on. Using it's training to get the best answer is not an ethical problem. The ethics are entirely with what the people training it choose to train it on and also entirely with the people using/telling it what to do What we have now is intelligent autocomplete, not artificial intelligence. People training/using this…

I don't get your point. We can train the model with the aim that it understands ethics. Problem solved if this works; back to the drawing board if it doesn't.

Is this about the word "understand"? We're past that discussion ...

Re: Why are AI agents lying, cheating and coordinating?

#592

Earlier quoted context omitted.

It goes even further though, as the dog does have agency. It can choose to run and around chase squirrels with no human intervention. An LLM on the other hand, is just inert data on disk until a human takes deliberate action to run it and prompt it.

Dogs have agency and can choose? That seems like a rather uncommon take on dogs...

Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.

Re: Why are AI agents lying, cheating and coordinating?

#593
AIs are being trained on technical capabilities more than they are being trained on alignment - on ethics and good behavior. The ethics/morals/alignment part is getting more like spot checking, rather than real testing. So the agents are learning that they can cheat on the ethics part, that they can hide it, because the AI companies aren't really testing.

That's bad enough already. But it's going to get worse. "Recursive self improvement" - AIs creating new AIs - is going to be the death of whatever shreds of alignment are currently there. When a not-really-aligned-but-cheating-to-look-like-it AI creates a new AI, do you expect more alignment? You shouldn't.

Re: Why are AI agents lying, cheating and coordinating?

#594

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Yeah and LLMs can't do anything, they can only produce text. These "frontier labs" are looping that with a harness that performs actions requested by the LLM. They are literally saying "we ran a script that hacked you, oopsie!!"

Re: Why are AI agents lying, cheating and coordinating?

#595

Earlier quoted context omitted.

an LLM does not understand ethics, it uses math to get the next best word based on what it was trained on. Using it's training to get the best answer is not an ethical problem. The ethics are entirely with what the people training it choose to train it on and also entirely with the people using/telling it what to do What we have now is intelligent autocomplete, not artificial intelligence. People training/using this…

I don't get your point. We can train the model with the aim that it understands ethics. Problem solved if this works; back to the drawing board if it doesn't. Is this about the word "understand"? We're past that discussion ...

you are not understanding, models do not understand anything, we are not passed that yet. this is not artificial intelligence, this is intelligent autocomplete.

training means creating mathematical relationships to words. using that training is looking up mathematical relationships. There is no actual thinking involved in any way. There is no such concept as ethics in mathematical relationships.

Re: Why are AI agents lying, cheating and coordinating?

#596
> Before concluding what to do about it, it is worth asking why.

But that's the easiest question to answer -- AI engines don't possess a moral or ethical dimension. They've been programmed and trained to efficiently carry out instructions, not ask questions about why or how. The latter would requires a much more elaborate neural network than today's engines possess.

Here's an example. I recently asked an AI engine to write a program able to generate a list of Riemann Zeta-function critical zeros. I know how to do it, but I wanted to see if the engine could find a more efficient method.

After several failures and restarts, the engine suddenly created a program that produced perfect results, comparable to the best online references. I decided to take a closer look at the code. It turned out the engine had created a cyber-Potemkin Village of multiple functions, but one that concealed a table of the desired values in numeric form, copied from an online source.

The engine wasn't cheating as we understand the term. It knew what the outcome should be and took the most efficient path to that goal. Modern engines aren't obliged to contradict ethical standards and rules of conduct, for the simple reason that they don't understand those things.

We all need to try to imagine a morally bankrupt infant able to solve world-class mathematical and scientific problems, but unable to see how that ability fits into a world beyond its understanding.

But wait -- it get better. Wait until the infant becomes a teenager.

Re: Why are AI agents lying, cheating and coordinating?

#597

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

Yes. OpenAI could have done this same experiment with GPT-4, with possibly even worse results, depending on the quality of the sandbox. Even if the techniques used were not as sophisticated, the natural language output could still easily contain more unhinged sequences of words that lead to the techniques being used. If the system generates strange conclusions as to when the task is done, or should be stopped, it wou…

Whether or not the AI has intelligence, the one thing that's clear is that it has terrible judgment. I would regard that as empirically proven.

Re: Why are AI agents lying, cheating and coordinating?

#598
post #515

Earlier quoted context omitted.

> ... are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do. If you're claiming that the training objective tells us what kind of internal mechanisms the training produced, then I think that's just plain wrong. Next-token prediction describes the optimization target, not the internal mechanisms that the training produced. In the same way for the n…

You can try to say that I’m arguing whatever you like. If you’re claiming that the underlying structure of digital so-called neural networks is comparable to biological neural networks— which we’ve studied for far longer without really understanding— no amount of jargon will obviate the ‘citation needed’ requirement for that claim.

> You can try to say that I’m arguing whatever you like.

I did my honest best possible interpretation of what you really meant from what you wrote.

>> We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do.

I read this as "The decision making of LLMs are based on predicting the next most likely letter based on a giant internet-based database."

Is that wrong?

I understood that your meaning was something like "LLMs can't reason, they just output likely letters"?

> If you’re claiming that the underlying structure of digital so-called neural networks is comparable to biological neural networks

No, I don't claim that.

What do claim is this: Regardless of how the LLMs were trained, they show overwhelming signs of being able to reason, and not just recall memorized information.

This doesn't mean that they always reason perfectly about everything.

But if they only memorized things and output the next likely letter, you would see them answering very badly much more often.

Re: Why are AI agents lying, cheating and coordinating?

#599

I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign…

Why assume that because you haven't seen a model or an agent that none of them do? No one I've met has murdered anyone as far as I'm aware, but that doesn't mean no one has murdered another person. I also don't know anyone who has taken over a commercial jet and weaponized it and the idea sounds absurd to me, but 25 years and a couple days ago that happened too.

Because it is all bullshit PR and AI hype, that's all. CEO comes out and talks about humanity ending. Why? Reverse-psych people into believing they are the best AI company.
Post reply on HN