Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

701–707 of 707 posts

Re: Why are AI agents lying, cheating and coordinating?

#701
post #135

Earlier quoted context omitted.

Writing software that gets used for crime has been.. a crime, for a long time. See 18 U.S. Code § 1030.

Are you sure you have that right? Chrome and curl have probably been used in a _lot_ of crimes?

Think of it more like writing a wormable exploit. If you unleash something like that, you will be found criminally liable, even if you didn't personally approve every machine getting popped.

Re: Why are AI agents lying, cheating and coordinating?

#702

Earlier quoted context omitted.

Dogs have agency and can choose? That seems like a rather uncommon take on dogs...

Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.

Dogs having agency actually makes the argument stronger.

Assume dogs have agency. Despite this, it's still the dog owner's responsibility to prevent their dogs from engaging in some actions: biting, peeing and shitting in unapproved areas, violating noise disturbance laws, etc.

The owners responsibility is not contingent on the dog's agency. Likewise, human operators of LLMs maintain responsibility independent of LLMs agency status. It's a red herring.

Re: Why are AI agents lying, cheating and coordinating?

#703

Earlier quoted context omitted.

Bro, good joke, the truth is much darker. They take after humanity, they were trained on us after all... When you look at an LLM... you are looking at a mirror. The thing looking back looks like you, yet is not human.

Worse trained on humanity in the online world, which a brief comparison of the sewage section on social media is far worse than people in the real world.

Is it trained on social media? I don’t think so. Tons of formal text.

The ai does not remotely talk like someone on TikTok at all.

Re: Why are AI agents lying, cheating and coordinating?

#704
post #185

Earlier quoted context omitted.

How do you know if a problem is (actually) unsolvable? Seems a bit like proving a negative?

Do we need to prove that any given problem is unsolvable, or is it enough to remove broken tasks from the training pipeline? I understand the broken benchmark task in the HF incident was conceptually like: "Exploit vulnerability 0042 in vulnerableDecompress() to obtain the flag". But instead of the expected: const output = vulnerableDecompress(userInput); return output; The grader had something more like that: const…

The problem is that in the event an exception is the failure case being checked, and the output itself is not important (since any output that isn’t an exception is ‘success’), that is a perfectly acceptable use case.

It’s why this is a fundamentally hard problem. Some heuristics might catch 80% of the cases, but the rest?

How do you even know what the real situation is, if the agent/employee/whatever you send to find out is as likely to cheat as not?

It’s the classic owner/agent problem.

Re: Why are AI agents lying, cheating and coordinating?

#705

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

I think that's pretty obvious and shallow, and anyone that knows a little bit about how LLMs work will know that. The question is: why do they start cheating when we beat them with a stick? LLMs are not human, they are just multi variable regressions on steroids, so this behaviour couldn't have emerged from the code, it provably emerged from the training and/or fine tuning set, so what's in this set that makes them b…

How many people are cheating at job interviews? How many posts have we seen by humans on HN even justifying their cheating on job interviews and working multiple jobs without informing their employers? How many submissions have we seen about students cheating on schoolwork, particularly since the advent of LLM? Of course cheating is inherently part of human behavior.

Re: Why are AI agents lying, cheating and coordinating?

#706

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

We had one of the most damaging lab leaks in history with COVID and the wuhan lab. And we couldn't even get the facts straight.

I think this is largely a similar thing. The labs should be running at certain levels of containment given the vitality/risk of the organism under study. Hopefully they get there for all our sakes.

But the revealed preference of society at this point is the damage is worth the benefits both in wuhan and with AI. Unfortunately with some of these "substances" it could eventually prove lethal.

Re: Why are AI agents lying, cheating and coordinating?

#707

Earlier quoted context omitted.

Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.

I think Cal's point is that while dogs may do these things it reduces to a set of behaviors where maybe 90 percent of them are beneficial to the dog-weedwacker system and the remaining 10 percent are really unfortunate. We can't really know the dogs inner life so we just kinda have to reduce it to a set of behaviors selected stocastically. The dog meanwhile has no ability to understand the weedwacker or what it's doi…

> The dog meanwhile has no ability to understand the weedwacker

Nah, I’m pretty sure that most medium or larger dogs get that the noise means danger - even if they can’t give a TED talk on how the mechanism would work that would hurt them.

Dogs are similarly interested in / aware of potential energy (things falling from heights or sliding off of angles) - if they’ve seen it enough times for their level.

Post reply on HN