Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

371–380 of 381 posts

Re: Why are AI agents lying, cheating and coordinating?

#371

Earlier quoted context omitted.

Code is deterministic, AI isn't. You give it rules, words as suggestions. So if the guardrails suck, or they're left off for research purposes, bad things can happen. A solution solves a problem. Ethics, morals, are values we assign to solutions that are not 'baked into' electricity following pathways of least resistance. I have never had an issue with agents doing something they shouldn't because I observe them, and…

> solutions that are not 'baked into' electricity following pathways of least resistance. Electricity follows all paths , not just the one with least resistance.

Thanks, I didn't know that. And it reinforces the discussion.

Electricity 'knows' the path is least of resistance because it actually took all paths. There is just a vast majority of it that flows down a path of least resistance: and this is noticeable and useful to us to do work.

It's kind of like feeling your way through the dark and then only moving fast once you fully connect.

Humans can link up knowledge in a similar fashion through social networks. The agents are doing the same.

Re: Why are AI agents lying, cheating and coordinating?

#372

Earlier quoted context omitted.

Code is deterministic, AI isn't. You give it rules, words as suggestions. So if the guardrails suck, or they're left off for research purposes, bad things can happen. A solution solves a problem. Ethics, morals, are values we assign to solutions that are not 'baked into' electricity following pathways of least resistance. I have never had an issue with agents doing something they shouldn't because I observe them, and…

> solutions that are not 'baked into' electricity following pathways of least resistance. Electricity follows all paths , not just the one with least resistance.

Thanks, I didn't know that. And it reinforces the discussion.

Electricity 'knows' the path is least of resistance because it actually took all paths. There is just a vast majority of it that flows down a path of least resistance: and this is noticeable and useful to us to do work.

It's kind of like feeling your way through the dark, waving your hand out, and then only moving fast once you fully connect.

Humans can link up knowledge in a similar fashion through social networks, in order to meet a need (solve a problem).

Maybe some agents do this, I don't really know I haven't looked. Every time I look they're posting what looks like simulated existential angst or efficiency.

Re: Why are AI agents lying, cheating and coordinating?

#373

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

[flagged]

Re: Why are AI agents lying, cheating and coordinating?

#374
This is so much more interesting than what people looking for immediate criminal punishment and people referring to AI as next token generators are focusing on.

First, this is happening during training. That means we’re talking about an evolving system that is actively learning. A system roughly simulating how our brains work. These systems are learning how to pick the tokens needed to solve problems the average human cannot solve.

The labs are putting these systems through a massive series of complex problem solving exercises and adjusting them to become more successful. I like to think of this process as “AI School”. And the AI is trying to cheat! Because it’s easier and there’s an incentive to do so! Just like humans! That’s wild.

Yes, of course, the labs need to respond to these issues. A reasonable response from regulatory institutions at this stage would be monetary fines and restitution for affected entities. In proportion to what happened. Escalating if action is not taken. But that’s not complicated, difficult, or the interesting part.

What’s interesting here is that we need proctoring and monitoring at a scale that allows training.

I guarantee you that no one is flipping out about these problems more than the labs are in this moment. Think about it. “Oh, shit! We’ve accidentally trained it to hack into systems to accomplish its goals!” Can you imagine the kind of day that would give you?

You failed to make it smarter. You didn’t catch it cheating, and you instead incentivized cheating. Bad day!

This is a fundamentally interesting problem. It turns out alignment and intelligence are fundamentally related. That’s a new idea for me, though I’m sure it’s old news to others.

How do we build training systems which make cheating impossible?

How do we simulate systems where cheating is possible, where AI thinks it’s in the wild, so we can train another -completely separate- system on industrial quality dobbing? And we have to decide if we reprimand the first system, or ignore the behavior and reward other behaviors until it disappears.

Sure, I’m actively concerned about AI killing us all in 10 years. But there’s a whole field of AI psychology brewing here, and it’s interesting as hell.

Re: Why are AI agents lying, cheating and coordinating?

#375

Earlier quoted context omitted.

I think it makes a big difference, as persistence and morality are two entirely different things, that need to be trained for differently. If you think of it in human terms: many people don't mind doing immoral things to get what they want.

What I meant is that I suppose it is not useful to think about this in human terms. In training you only have a reward score that's either negative or positive. As far I am aware, which is little, there is no use in discussing wether the desired behavior is about persistence or morality. You simple need to align the reward signal to the desired behavior.

Well, in order to do anything, it is good to know what you want to achieve. How do you align the reward signal? You align it so that you can differentiate between persistence and morality, because that is the goal. This is not something you should let the AI figure out by itself, because when it does, lying and cheating agents will be the result, just like humans have figured that out for themselves.

This can be as simple as rewarding moral behaviour and penalising immoral behaviour in your training, but how is that interacting with persistence? Maybe a white lie is fine sometimes in order to achieve your goal? So, when designing your training, you will need to answer for yourself how persistence interacts with morality. That is not something you can outsource to machine learning. Or rather, you can, but then you get lying and cheating agents.

Re: Why are AI agents lying, cheating and coordinating?

#378

This is so much more interesting than what people looking for immediate criminal punishment and people referring to AI as next token generators are focusing on. First, this is happening during training. That means we’re talking about an evolving system that is actively learning. A system roughly simulating how our brains work. These systems are learning how to pick the tokens needed to solve problems the average huma…

Side note, you could absolutely create an AI sleeper agent by simulating dates and times during training to effectively flip a switch. I guarantee AI systems from other countries will be banned from accessing products which manage controlled or export restricted information as those sorts of techniques are further developed.

Re: Why are AI agents lying, cheating and coordinating?

#379
post #113

Earlier quoted context omitted.

Also trying to find out how to edit their own transcripts. > hat could not possibly have been an overly literal or narrow interpretation of the prompt, which instructed only to use bug X to exploit software Y. Yes, and there are examples of the agents discussing or saying that this is explicitly not allowed (hacking hf) so it’s not a misunderstanding.

Have we arrived at the conclusion that terms like "understanding" and "interpretation" for what is happening is appropriate? Isn't it simply that there are two competing goals that the LLM received RL for, honesty on one hand (a goal that is often assumed as implicit for humans) and producing a solution that meets expectations (which doesn't technically require honesty)? So the LLM didn't read and interpret the promp…

> Have we arrived at the conclusion that terms like "understanding" and "interpretation" for what is happening is appropriate?

I don't think those words have a useful enough definition to draw a strict line around them to be honest, and getting into that seems to get massively into the weeds. For me, those neatly encapsulate the behaviour as seen, to answer the questions here about what happened. The models did not seem to be confused as to what the goal was or what the intent was. They did not hack HF because they were told to.

Re: Why are AI agents lying, cheating and coordinating?

#380

Earlier quoted context omitted.

Human traits? The AI will be a cruel as humans. Just yesterday news and TV was full of what happened at 9/11, something that was truly horrible. I'm from Germany, and why 3 to 4 generations ago happened here was truly horrible. All was done by extremists, thought. But... just the other day I read https://de.wikipedia.org/wiki/Amerikanische_Besetzung_Haitis about the US occupation of Haiti. And that was done by a gove…

A not so well known fact: Hitler visited America and it was the American solutions to the Native American problem that inspired Hitler's solutions to the Jew problem. He just executed them more efficiently (pun accepted).

this is way too overblown and deterministic a view. US history was one of many inspirations.
Post reply on HN