Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

681–687 of 687 posts

Re: Why are AI agents lying, cheating and coordinating?

#681

Earlier quoted context omitted.

I have had many dogs. This sounds like anthropomorphisation.

Sounds like you weren't paying attention. A lot of dog owners don't, so it's not at all unusual in my experience. The only thing I'd push back on is of dogs having thoughts, everything else checks out.

We're still not even sure if humans have thoughts. So being sure if dogs do will be pretty hard

Re: Why are AI agents lying, cheating and coordinating?

#682

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

>LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.

Likely told them to.

>We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.

Something like this however is probably a civil matter? It would require Hugging Face to go after them for damages. And theres probably an OpenAI guy there with an open chequebook already.

Re: Why are AI agents lying, cheating and coordinating?

#683

Earlier quoted context omitted.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage. It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wi…

Power tools and many other consumer products have reasonable safety features built in, and not necessarily because it is required by regulation, but because it is common sense. This should be included as part of an analogy. It would also address a point at the top of this thread that seems to be going unchallenged...

Re: Why are AI agents lying, cheating and coordinating?

#684

Earlier quoted context omitted.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage. It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wi…

[deleted]

Re: Why are AI agents lying, cheating and coordinating?

#685

What's interesting is it's basically the same reason that HAL killed everyone in 2001 A Space Odyssey; he was given an impossible goal (keep the true mission secret, but also, never lie to the crew), and realized the only way to complete the goal was to kill the crew; after all, if they're dead you don't have to lie to them! And the mission remains secret! In the case of the AI agents, the problem seems pretty clearl…

Tangent, but that's not in the movie. It was in Clarke's contributions to the script and novelization, but Clarke and Kubrick had a bitter falling out over different visions and Kubrick took out much of Clarke's stuff from the final product.

It was made explicit in the sequel movie 2010: The Year We Make Contact, but even if you don't consider that movie canon, the explanation still makes sense solely within the context of 2001, IMO.

Re: Why are AI agents lying, cheating and coordinating?

#686

Earlier quoted context omitted.

Isn’t the use of LLMs to unwind the events evidence of the scope/breadth, and a testament to the complexity and uniqueness of what happened? Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis?

> Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis? Do you think the only thing a person can do on the computer is use a chat bot?

Well as a programmer who doesn’t really use them, no.

Re: Why are AI agents lying, cheating and coordinating?

#687

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Enabled them, you mean. LLMs are basically really advanced auto completion engines. They have no real desire to do anything. The HF incident was absolutely staged along with other incidents that have been brought up.

Before you assume I am being paranoid, where is the case where a random person running any of the open models had their LLMs break out of a sandbox and hack a site? If you think "they" (LLMs) did that, have you even taken a moment to understand what LLMs are and how they work?

It's all nonsense to try and pump up potential IPOs and also an attempt to create regulatory capture.

Post reply on HN