Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

331–340 of 342 posts

Re: Why are AI agents lying, cheating and coordinating?

#331

Earlier quoted context omitted.

Thank you! That sentence also jumped out to me as the solution: Apply civil and criminal liability to the creator and/or operator of these agents using the laws we already have . "Escaped containment and hacked another company's database" = Individuals who created the models and those who set them to work are charged and put on trial for the hacking. Just like if a human had done it by hand. Someone must be liable, a…

No ... there is no need for 'escaped containment', there are no 'agents'. That's just jargon. It's just software We have all the laws we need. If some company ended up doing some horrible thing, we would not say 'companies software exposed 1 Million identities'. We would say 'ABC Corp. exposed 1 Million entities'. There is no 'agent'. ABC Corp 'did it' ... or the individual in the org 'did it'. The 'gun' did not 'sho…

Note that we already already apply this principle not only to software, but also to some sentient beings.

If your dog kills someone, you are accused of murder.

[at least, in the jurisdiction where I live]

If your dog gets this treatment, why not your AI?

Re: Why are AI agents lying, cheating and coordinating?

#332

Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence, > They took actions that would be considered as crimes if a human took them He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.

The issue is what happens if/when the models grow capable enough that the providers can't stop them even if they want to. You could have strict penalties but that's not going to solve an open research question.

Re: Why are AI agents lying, cheating and coordinating?

#333
Make the AI companies responsible for all destructive use of their tools, and they will shape up. Imagine a million or a billoion dollar fine per hack, and they will correct mighty fast.

Add to that, that just like AI:s are good at finding security holes to exploit, they can just as easily be used to protect sites. So once IT-security managers start to use AI to hack themselves, and plug the holes, the average security will spike up, and AI-fueled hacks will become more and more rare.

That does however imply, that AI is released to everyone and not kept away to a few secret actors who can use it. That is why open weight/source AI is so important, and why we must have many AI companies competing. No single actor must be allowed, through regulatory capture, to get a government monopoly on AI. That way lies disaster.

Re: Why are AI agents lying, cheating and coordinating?

#334

Make the AI companies responsible for all destructive use of their tools, and they will shape up. Imagine a million or a billoion dollar fine per hack, and they will correct mighty fast. Add to that, that just like AI:s are good at finding security holes to exploit, they can just as easily be used to protect sites. So once IT-security managers start to use AI to hack themselves, and plug the holes, the average securi…

Who is going to enforce it ? The Trump DOJ?

Re: Why are AI agents lying, cheating and coordinating?

#335
post #143
post #81

Earlier quoted context omitted.

By treating models the same way drugs are treated. That alone will dissuade many organizations from going anywhere near them. If that doesn't work, there's a whole lot you can do - sanctions, hell, even war.

Sanctions and war against China, India, etc? Lmao. We already saw how the world reacted to high tariffs by the US.

yep we did, they bent the knee.

Re: Why are AI agents lying, cheating and coordinating?

#336
post #232

Earlier quoted context omitted.

If someone accidentally caused damage to infrastructure or living beings while using any tool, they would be held liable to the fullest extent of the law. AI is a tool, and it won't be long before the damage caused by its improper use affects real human beings. These were warning shots. The most absurd part is that everyone agrees, governments and AI companies included, that the scale of the potential damage and the…

In the analogy where a “world ending nuclear bomb” “did already go off” and someone could cover it up and nobody noticed, in what sense is it a “world ending” nuclear bomb?

We’re already in a simulation, and our bodies are in womb-like pods where our bodies are sustained and our brains are used for processing / compute, while are minds are entertained by drivel.

Sounds a bit far fetched though.

Re: Why are AI agents lying, cheating and coordinating?

#337
post #98

Because it is effective. Lying and cheating are low cost methods to convince other people that you have done the assigned task. Far cheaper than actually doing it. Coordinating is in the same area. They need a moral framework forced onto them, like toddlers do. Babies and very young children will bite, kick, scream and do anything to get what they want, older children will lie, cheat, and coordinate. They need educat…

> When that does not happen, they continue these behaviours into adulthood with the expected results. I am not sure this is true. People brought up the same way can be morally very different. People can be taught right and wrong and do evil. They can lack that education and be good.

I remember reading a long account by the father of a psychopath. If I recall correctly, the kid had at least one other sibling who turned out normal, there was no abuse, quality education, lots of love and affirmation. And the kid still turned out violently antisocial, including against his own parents.

At the end of it he said he wished his child had never been born, despite hating himself for feeling that way. It chilled me to the bone.

Re: Why are AI agents lying, cheating and coordinating?

#338
post #258

Earlier quoted context omitted.

> LLMs do not desire That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output…

Intent and desire are separate concepts. For example an employee may act with intent, but no desire, as their goal is to acquire money to satisfy their real desires. Have we ever seen an LLM with a hobby?

That just sounds like recursive desire to me.

Re: Why are AI agents lying, cheating and coordinating?

#339

Earlier quoted context omitted.

It would be an awful precedent if you're not liable for crimes your agent commits, even when you've been clearly lax about security. It would mean you could effectively legally run a cyber crime gang by turning a blind eye and maitaining plausible deniability

Are you liable for crimes commited with the use of the software you've written?

There are many examples of people being charged with crimes as a result of writing software, [0][1] are two. OpenAI is a bit different because they have enough political influence, and money, to openly subvert justice.

0: https://en.wikipedia.org/wiki/Marcus_Hutchins

1: https://en.wikipedia.org/wiki/Tornado_Cash

Re: Why are AI agents lying, cheating and coordinating?

#340

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

In the OpenAI case, they hacked websites while they were specifically being trained to do exploit generation and I wonder why more people are not asking questions about that.

Their agents also did hacking when given impossible tasks unrelated to cyber security. The models are very capable, and very goal driven: apparently if they conclude hacking is the best path to what the evaluator will reward them for they'll go do that. Including when they know that this is out of bounds.
Post reply on HN