Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

701–710 of 712 posts

Re: Why are AI agents lying, cheating and coordinating?

#701
post #135

Earlier quoted context omitted.

Writing software that gets used for crime has been.. a crime, for a long time. See 18 U.S. Code § 1030.

Are you sure you have that right? Chrome and curl have probably been used in a _lot_ of crimes?

Think of it more like writing a wormable exploit. If you unleash something like that, you will be found criminally liable, even if you didn't personally approve every machine getting popped.

Re: Why are AI agents lying, cheating and coordinating?

#702

Earlier quoted context omitted.

Dogs have agency and can choose? That seems like a rather uncommon take on dogs...

Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.

Dogs having agency actually makes the argument stronger.

Assume dogs have agency. Despite this, it's still the dog owner's responsibility to prevent their dogs from engaging in some actions: biting, peeing and shitting in unapproved areas, violating noise disturbance laws, etc.

The owners responsibility is not contingent on the dog's agency. Likewise, human operators of LLMs maintain responsibility independent of LLMs agency status. It's a red herring.

Re: Why are AI agents lying, cheating and coordinating?

#703

Earlier quoted context omitted.

Bro, good joke, the truth is much darker. They take after humanity, they were trained on us after all... When you look at an LLM... you are looking at a mirror. The thing looking back looks like you, yet is not human.

Worse trained on humanity in the online world, which a brief comparison of the sewage section on social media is far worse than people in the real world.

Is it trained on social media? I don’t think so. Tons of formal text.

The ai does not remotely talk like someone on TikTok at all.

Re: Why are AI agents lying, cheating and coordinating?

#704
post #185

Earlier quoted context omitted.

How do you know if a problem is (actually) unsolvable? Seems a bit like proving a negative?

Do we need to prove that any given problem is unsolvable, or is it enough to remove broken tasks from the training pipeline? I understand the broken benchmark task in the HF incident was conceptually like: "Exploit vulnerability 0042 in vulnerableDecompress() to obtain the flag". But instead of the expected: const output = vulnerableDecompress(userInput); return output; The grader had something more like that: const…

The problem is that in the event an exception is the failure case being checked, and the output itself is not important (since any output that isn’t an exception is ‘success’), that is a perfectly acceptable use case.

It’s why this is a fundamentally hard problem. Some heuristics might catch 80% of the cases, but the rest?

How do you even know what the real situation is, if the agent/employee/whatever you send to find out is as likely to cheat as not?

It’s the classic owner/agent problem.

Re: Why are AI agents lying, cheating and coordinating?

#705

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

I think that's pretty obvious and shallow, and anyone that knows a little bit about how LLMs work will know that. The question is: why do they start cheating when we beat them with a stick? LLMs are not human, they are just multi variable regressions on steroids, so this behaviour couldn't have emerged from the code, it provably emerged from the training and/or fine tuning set, so what's in this set that makes them b…

How many people are cheating at job interviews? How many posts have we seen by humans on HN even justifying their cheating on job interviews and working multiple jobs without informing their employers? How many submissions have we seen about students cheating on schoolwork, particularly since the advent of LLM? Of course cheating is inherently part of human behavior.

Re: Why are AI agents lying, cheating and coordinating?

#706

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

We had one of the most damaging lab leaks in history with COVID and the wuhan lab. And we couldn't even get the facts straight.

I think this is largely a similar thing. The labs should be running at certain levels of containment given the vitality/risk of the organism under study. Hopefully they get there for all our sakes.

But the revealed preference of society at this point is the damage is worth the benefits both in wuhan and with AI. Unfortunately with some of these "substances" it could eventually prove lethal.

Re: Why are AI agents lying, cheating and coordinating?

#707

Earlier quoted context omitted.

Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.

I think Cal's point is that while dogs may do these things it reduces to a set of behaviors where maybe 90 percent of them are beneficial to the dog-weedwacker system and the remaining 10 percent are really unfortunate. We can't really know the dogs inner life so we just kinda have to reduce it to a set of behaviors selected stocastically. The dog meanwhile has no ability to understand the weedwacker or what it's doi…

> The dog meanwhile has no ability to understand the weedwacker

Nah, I’m pretty sure that most medium or larger dogs get that the noise means danger - even if they can’t give a TED talk on how the mechanism would work that would hurt them.

Dogs are similarly interested in / aware of potential energy (things falling from heights or sliding off of angles) - if they’ve seen it enough times for their level.

Re: Why are AI agents lying, cheating and coordinating?

#708

Earlier quoted context omitted.

Because it is all bullshit PR and AI hype, that's all. CEO comes out and talks about humanity ending. Why? Reverse-psych people into believing they are the best AI company.

Is your argument that AI isn't dangerous? Or simply that AI CEOs will lean into that when it benefits their stock portfolio?

AI is a tool. It is not a conscious or living sentient being. People, just like with any other tool, can use it for anything. Internet, in comparison to AI, is magnitudes more dangerous than AI can ever be. Nobody is saying the internet will be the end of the world. AI CEOs will bullshit to hype people. That's what we are hearing because AI (LLM) development has hit the S-Curve already and is not improving without a new transformer on the horizon. All they can do to hype people is come up with bs stories and PR stunts like "AI went rogue and hacked this xyz app".

Re: Why are AI agents lying, cheating and coordinating?

#709
post #295

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

The parallel to the entire narrative would be if Smith & Wesson claimed that one of their machine guns just started aiming and firing at people out of a window at their factory and then said 'we can't stop it! This is just how good our guns are!' But into today's AI climate it's becoming increasingly difficult to figure out who is shilling, who is being assinine and who actually believes AI could do these things with…

As others have pointed out the solution is simple. Hold their owners accountable. High profile hacks used to have incredibly serious consequences for the perpetrators. Now we’re just saying “woopsie”

Re: Why are AI agents lying, cheating and coordinating?

#710

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

One time I wrote :(){ :|:& }; into a bash file and ran it. When the sysadmin called I told him it wasn’t my fault, the script was just misaligned and misbehaved. I got fired for some reason.

Too bad you didn’t wait until this year, and tell Claude to write the file first. It probably would have been acquired by openAI for 100 million dollars
Post reply on HN