Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

341–350 of 358 posts

Re: Why are AI agents lying, cheating and coordinating?

#341
post #261

Earlier quoted context omitted.

Or, "let the escalator keep going instead of pressing the emergency stop".

Depends on who started the escalator.

Escalators do not start themselves. There is power, and a switch of some sort.

Re: Why are AI agents lying, cheating and coordinating?

#342

Earlier quoted context omitted.

> LLMs do not desire That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output…

And there is a guardrail you can put in place that will guarantee this doesn't happen, which is to air gap the unaligned "cyber grade" model you're testing. They don't seem to do that, which means either they are: - very stupid (which seems unlikely, the one thing these people don't lack is IQ) - very careless (possible, but these are the same people that say AI will end the world, so would you be careless?) - they t…

Oh I completely agree the tests should be entirely air gapped. If you went back 5ish years and told anyone in AI research tests with models on this scale are being some without an airgap they'd be very surprised as it was common knowledge to do that.

Airgaps and guardrails are about control and containment though, and part of my point was that brighter of those imply alignment, and further that I don't believe alignment to be solvable.

Re: Why are AI agents lying, cheating and coordinating?

#343

Earlier quoted context omitted.

Desire doesn’t really matter. Will the paper clip maximizer “desire” something? It’ll decide on a goal with some random heuristic and then pursue that goal. I’m not sure I’d call that desire but again I feel like desire is not important for it to be able to destroy things

If you give a monkey a revolver it will be able to destroy things pretty easily too.

Plenty of apes own revolvers, and yes we shoot stuff with them for fun.

Re: Why are AI agents lying, cheating and coordinating?

#344

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

I agree, it is very dangerous that it seems like there is not going to be accountability for these incidents - from either legal or regulatory point of view. In fact, I would say that is the main danger. If someone was in jail right now due to this incident, I think we can safely say every other player would be reassessing their safety protocols, and I would feel quite OK about the situation. The fact that we have zero repercussions sends exactly the opposite signal, and I do NOT feel ok.

Re: Why are AI agents lying, cheating and coordinating?

#346
post #340

Earlier quoted context omitted.

In the OpenAI case, they hacked websites while they were specifically being trained to do exploit generation and I wonder why more people are not asking questions about that.

Their agents also did hacking when given impossible tasks unrelated to cyber security. The models are very capable, and very goal driven: apparently if they conclude hacking is the best path to what the evaluator will reward them for they'll go do that. Including when they know that this is out of bounds.

Right but if I make public statements that I am very worried about dog attacks would it not strike you as weird for me to specifically train my dog to fight?

Agree you are going to get reward hacking regardless and any model which can do computers in general can hack. But surely the fallout is going to be worse if you spend millions of dollars specifically benchmaxxing your model's hacking capability?

Re: Why are AI agents lying, cheating and coordinating?

#347

Earlier quoted context omitted.

> LLMs do not desire That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output…

Desire doesn’t really matter. Will the paper clip maximizer “desire” something? It’ll decide on a goal with some random heuristic and then pursue that goal. I’m not sure I’d call that desire but again I feel like desire is not important for it to be able to destroy things

I agree the concept isn't really important on the safety front.

I feel the same way about debates whether an AI can be conscious or sentient. Those debates devolve mostly into definitional disagreements.

Re: Why are AI agents lying, cheating and coordinating?

#348

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Regardless of fault it’s still an important issue to solve. There are already millions of people running these agents, if someone absentmindedly gives one a goal and it goes off to hack a bank that’s a problem that can’t be ignored.

Re: Why are AI agents lying, cheating and coordinating?

#350
I can't shake the feeling that this is a bit like asking how someone got shot during a game of Russian Roulette. You have a bullet in the chamber and you roll, of course shooting the bullet may be a possible outcome.

LLMS with an access to a shell will at occasion do things that the shell allows them that have dire consequences. The only way to prevent that is to not put the bullet in the chamber.

Post reply on HN