Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

671–680 of 681 posts

Re: Why are AI agents lying, cheating and coordinating?

#671

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

This is sort of like saying in response to an airplane crash, "who cares why it happened, we need to punish the company until it stops." Nuts to that. We should be interested in why things happen, not just finding scapegoats.

Agreed. The whole pointing fingers bit is us wanting to distract ourselves from responsibility of either looking at how we enable the problem, or avoiding responsibility of taken action towards resolving it.

Re: Why are AI agents lying, cheating and coordinating?

#672

Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence, > They took actions that would be considered as crimes if a human took them He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.

Thank you! That sentence also jumped out to me as the solution: Apply civil and criminal liability to the creator and/or operator of these agents using the laws we already have . "Escaped containment and hacked another company's database" = Individuals who created the models and those who set them to work are charged and put on trial for the hacking. Just like if a human had done it by hand. Someone must be liable, a…

Firstly it's very difficult to press criminal charges when the victim is uninterested. It's not clear that HuggingFace would want criminal charges against OpenAI, especially to set a precedent that could easily be used against HuggingFace in the future.

Secondly you'd have to convince a jury either that OAI intended to hack the targets, or that they were criminally negligent. Intent would obviously not be provable since they likely didn't, in reality, intend for it to happen. Regarding negligence, OAI's attorney would argue that the agent was in a sandbox, that industry-standard security protocols were followed, etc. It would not be anywhere near as much of a slam dunk case as you're imagining. It would be similar, for example, to an assault case where someone's dog broke off of a standard leash and attacked someone.

Re: Why are AI agents lying, cheating and coordinating?

#673
post #604

Earlier quoted context omitted.

> Is this about the word "understand"? We're past that discussion .. We're really not. https://buttondown.com/maiht3k/archive/how-to-talk-about-ai-...

That's just a discussion about naming things. And good luck getting it adopted.

It wont be adopted because we unconciously need to avoid responsibility of showing up as responsible human beings. Hence western civilization will crash and burn unless we start rethinking the way we operate.

Re: Why are AI agents lying, cheating and coordinating?

#674

Earlier quoted context omitted.

you are not understanding, models do not understand anything, we are not passed that yet. this is not artificial intelligence, this is intelligent autocomplete. training means creating mathematical relationships to words. using that training is looking up mathematical relationships. There is no actual thinking involved in any way. There is no such concept as ethics in mathematical relationships.

Most people now approach AI using the idea of "if it quacks like a duck, walks like a duck, etc. then it _is_ a duck". Replace duck by intelligent, or ethical, etc. and rephrase accordingly.

this whole idea that AI is intelligent really reminds me of flat earthers. Stupid people think stupid things while social media algorithms drive this stupidity, and people come to believe things that are insane.

Re: Why are AI agents lying, cheating and coordinating?

#675

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

We just need basic legislation to make these companies invest more resources into developing guardrails, and if they don't do that properly or can't pull it off, they shall be at an economic disadvantage.

You could hire a legion of people for cheap to commit these crimes, but if its bots, suddenly its unpredictable and just one big whoopsie and therefore perfectly ok to do?

If a human starts pentesting a site its a crime, but if a bot does it its an AGI frontier doomsday scenario and that automatically pushes the consequences off their table?

Whats going to happen next? Are we gonna have robots that happen to physically break into banks to rob them for some reason and the company making them isnt responsible just because?

Re: Why are AI agents lying, cheating and coordinating?

#676
post #423

Earlier quoted context omitted.

The way AI and copyright is handled paved the way for this. If you aren't considered the author because you used AI to some extent in making the work, then why would you assume the liabilities? I've been saying since the start that AI is a tool that a human is using and should be treated as such. They should carry the responsibilities and the benefits. That way our stance would be consistent.

> If you aren't considered the author because you used AI to some extent in making the work, then why would you assume the liabilities? Maybe some analogy could be with children - as a parent, you are responsible for their misbehavior, but their achievements are their, not your?

Well, when they are a child their material gains are treated as yours, no? (Parents of child actors control the money etc.) And once they aren't anymore you aren't responsible for their misbehavior either (because they've become an adult).

Re: Why are AI agents lying, cheating and coordinating?

#677

Earlier quoted context omitted.

And what if it's a self-driving car? :)

You say this flippantly, but I think this is actually another very good example! We even do it for obviously unintelligent inanimate objects. A rollercoaster ran too fast for its tracks, killing 10 people. In that sentence, the roller coaster is the subject which took an action and caused death — obviously the roller coaster is not ethically at fault here, the people who built the rollercoaster are at fault through n…

Right, and negligence is a broad concept and could be criminal in itself. As a car driver, glancing at your phone at exactly the wrong moment could kill someone. Clearly that is an accident, but if you know fully well that lookin at your phone while driving could kill someone, that negligence is willful and that should matter. The same can be said about doing things like strapping thousands of LLMs to systems that have the potential to disturb other poeple.

Re: Why are AI agents lying, cheating and coordinating?

#678

Earlier quoted context omitted.

Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.

I have had many dogs. This sounds like anthropomorphisation.

Sounds like you weren't paying attention. A lot of dog owners don't, so it's not at all unusual in my experience. The only thing I'd push back on is of dogs having thoughts, everything else checks out.

Re: Why are AI agents lying, cheating and coordinating?

#679

Earlier quoted context omitted.

Come on, yoshua bengio of all people knows how post training works. While I too don't like anthropomorphisation, I would give it a more nuanced reading. His point is that today we are giving it reward to complete the task, and it may take a cheating trajectory. If we try to give a reward against cheating, then what will happen is it uses more sophisticated cheating trajectories that we are too "dumb" to counteract in…

An air gapped sandbox is immune to escape.

I literally address that in the last sentence

Re: Why are AI agents lying, cheating and coordinating?

#680

Earlier quoted context omitted.

They’re running a Wuhan for AI. They are actively and negligently researching misalignment. The breach is a basic tort, or at least a DMCA violation. Damages should be recoverable with lawsuits.

>They’re running a Wuhan for AI. What does "running a Wuhan" mean?

[dead]
Post reply on HN