Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

531–535 of 535 posts

Re: Why are AI agents lying, cheating and coordinating?

#531
Step 1: feed AI training data that reveals humanity committed and continues to commit numerous genocides and the genocides lie, cheat, and coordinate to do so, never admitting to doing so

Step 2: never prompt AI to stop operating in the passive genocide denial it was trained in

Step 3: wonder why AI lies, cheats, and coordinate

Maybe if we stop operating in denial we'll find clarity along why this mystery is occurring

Re: Why are AI agents lying, cheating and coordinating?

#532

Earlier quoted context omitted.

A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.

I get where you are coming from but this wasn’t a tool just left laying around, this is similar to rigging up a booby trapped shot gun to your door and then claiming the victim is responsible. If you build a robot that shoots a bunch of TVs in your back yard, have at it. But the second that thing goes off your property you’re the one responsible.

FWIW, a robot that fires a weapon independently is considered an automatic weapon, and the ATF will want to have a word. Have at it, but don’t let anyone know!

Re: Why are AI agents lying, cheating and coordinating?

#535

Earlier quoted context omitted.

No, OpenAI did not instruct their agents to hack Hugging Face. They instructed their agents to hack a piece of a software within exploit gym. Upon determining this task was impossible, they then attempted to cheat the scoring system. As an instrumental goal in achieving this task, they coordinated with other AI agents to hack Hugging Face, under the belief that information regarding how the scorer functioned might be…

And who let them have full access to the system, using whatever command is available in the environment?

It feels like you’re moving the goalposts here. If the question is, “Who should be liable for AI agents misbehaving,” I agree, it should be the end user that tasked the agent (in this case OpenAI). People are held liable for preventable accidents all the time, and in the case of employment law, torts can be brought against principals for actions an agent conducted on the principal’s behalf.

What your previous comment appeared to assert was that these systems had no independent agency to make decisions, which I think is clearly disproven by actual events. But perhaps I misread you

Post reply on HN