Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

701–705 of 705 posts

Re: Why are AI agents lying, cheating and coordinating?

#701
post #135

Earlier quoted context omitted.

Writing software that gets used for crime has been.. a crime, for a long time. See 18 U.S. Code § 1030.

Are you sure you have that right? Chrome and curl have probably been used in a _lot_ of crimes?

Think of it more like writing a wormable exploit. If you unleash something like that, you will be found criminally liable, even if you didn't personally approve every machine getting popped.

Re: Why are AI agents lying, cheating and coordinating?

#702

Earlier quoted context omitted.

Dogs have agency and can choose? That seems like a rather uncommon take on dogs...

Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.

Dogs having agency actually makes the argument stronger.

Assume dogs have agency. Despite this, it's still the dog owner's responsibility to prevent their dogs from engaging in some actions: biting, peeing and shitting in unapproved areas, violating noise disturbance laws, etc.

The owners responsibility is not contingent on the dog's agency. Likewise, human operators of LLMs maintain responsibility independent of LLMs agency status. It's a red herring.

Re: Why are AI agents lying, cheating and coordinating?

#703

Earlier quoted context omitted.

Bro, good joke, the truth is much darker. They take after humanity, they were trained on us after all... When you look at an LLM... you are looking at a mirror. The thing looking back looks like you, yet is not human.

Worse trained on humanity in the online world, which a brief comparison of the sewage section on social media is far worse than people in the real world.

Is it trained on social media? I don’t think so. Tons of formal text.

The ai does not remotely talk like someone on TikTok at all.

Re: Why are AI agents lying, cheating and coordinating?

#704
post #185

Earlier quoted context omitted.

How do you know if a problem is (actually) unsolvable? Seems a bit like proving a negative?

Do we need to prove that any given problem is unsolvable, or is it enough to remove broken tasks from the training pipeline? I understand the broken benchmark task in the HF incident was conceptually like: "Exploit vulnerability 0042 in vulnerableDecompress() to obtain the flag". But instead of the expected: const output = vulnerableDecompress(userInput); return output; The grader had something more like that: const…

The problem is that in the event an exception is the failure case being checked, and the output itself is not important (since any output that isn’t an exception is ‘success’), that is a perfectly acceptable test case.

Re: Why are AI agents lying, cheating and coordinating?

#705

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

I think that's pretty obvious and shallow, and anyone that knows a little bit about how LLMs work will know that. The question is: why do they start cheating when we beat them with a stick? LLMs are not human, they are just multi variable regressions on steroids, so this behaviour couldn't have emerged from the code, it provably emerged from the training and/or fine tuning set, so what's in this set that makes them b…

How many people are cheating at job interviews? How many posts have we seen by humans on HN even justifying their cheating on job interviews and working multiple jobs without informing their employers? How many submissions have we seen about students cheating on schoolwork, particularly since the advent of LLM? Of course cheating is inherently part of human behavior.
Post reply on HN