Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

551–560 of 564 posts

Re: Why are AI agents lying, cheating and coordinating?

#551
post #393
post #352

Earlier quoted context omitted.

I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. All that while still not knowing how either kind actually works.

I always wonder what makes people take the other side of this argument. They do it quite passionately. Why actively encourage viewing LLMs as human? Who is that benefitting?

Does the argument require benefit? Isn’t the argument based on caution?

I haven’t heard many people explicitly saying “these things behave like humans”, but more generally “we don’t even know how to define human consciousness, we don’t have a thorough grasp of how the brain works, we are still very much in the dark on a lot of these topics, so how can we say one way or the other?”

In other words, agnosticism: I don’t know.

In general, it’s baffling to me that anyone has an unshakable opinion on what exactly is happening. It seems like raw egotistical hubris.

Re: Why are AI agents lying, cheating and coordinating?

#552

Earlier quoted context omitted.

Can’t a prosecutor charge them regardless?

In Indian legal syatem a case can be filed suo moto by the judges or agencies. You don't require the affected party to sue. Not sure how it works in the US.

For civil suits, you typically need the affected parties to sue. Otherwise, who claims the damages?

But this isn’t just civil, it’s criminal. Hacking is a criminal offense. This could be a CFAA violation. That’s landed people life in prison before. There, you don’t need the victims to be motivated. The federal prosecutors could just go ahead.

Re: Why are AI agents lying, cheating and coordinating?

#553

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

Yes and no.

If your buddy leaves his car parked at the top of a hill without the parking brake on and it rolls down the hill and side-swipes a bunch of vehicles and narrowly misses an elderly person walking by with a cane someone could easily say:

"Dude wtf is wrong with you, you left your car parked on the top of a hill with no brake and let it roll into traffic"

The phrasing doesn't absolve the offender of their negligent behaviour and the consequences of it.

The only thing thing does is the lack of action from regulators and society writ large.

Our lack of action is what allows people like Sam Altman and Dario and the irresponsible people who choose to work for them to be continue to be negligent.

Re: Why are AI agents lying, cheating and coordinating?

#554

Earlier quoted context omitted.

Why is anthropomorphism the problem here? If OpenAI hired a contractor and they did this, OpenAI or the contractor would still be liable, depending on the contract language.

A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.

Right. Among bicycle advocacy groups it's been well known for long time that cars do not run over people, drivers do.

The fact that we talk about a car running someone over, and this is the same in many different languages and countries, contributes to lower punishments for drivers. Clearly it was just an accident. He or she was run over by a car.

Now we see that same language tricks play out again every time an LLM did something illegal.

Re: Why are AI agents lying, cheating and coordinating?

#555

Earlier quoted context omitted.

This is the problem with optimization generally, even in the human domain. You measure task performance with a metric and punish/reward based on the metric. Anyone who likes reward / hates punishment isn't going to actually care about doing the task well, they are going to care about the metric. The models know that we want them to do things, but also from the training corpus that we evaluate performance using benchm…

I agree with most of this, but you're misunderstanding "alignment" as coined. Yes, training powerful enough AI, any simple optimization target gets you malign behavior, because human values are not simple. If you insist on making powerful AI, you'd better instill respect for human values! That's "alignment". https://www.lesswrong.com/posts/ZxWzCGKzX84S7DBZ9/when-was-t...

How do you do that in the current paradigm other than creating yet another gameable metric? And something I didn't mention above is that there is no difference between "solving the task" and "optimizing the metric" for an ML model, even though there clearly is for us. So it's not clear to me how you "fix" something that is baked into the architecture. All I'm saying is "instilling respect for human values" is not something that can actually be done via a cost function. In no small part because we humans probably don't even agree on those values, let alone on a single metric with which to quantify and "optimize" them.

For example, we agree that "merit" is valuable and that we should reward "merit." But to reward it we have to quantify it, and what metric should we use? Raw SAT score to get into college? But that also captures socioeconomic factors that unfairly penalize some and reward others. We generally agree that those who provide more value should earn more money, but what does that look like? Do we all agree on what activities are or should be valuable, or on how they should be rewarded? Until recently, I thought we all agreed that "empathy" was a human value, but a lot of people in this space, who are making these decisions unilaterally for all of us, don't apparently share that belief.

Re: Why are AI agents lying, cheating and coordinating?

#556

Earlier quoted context omitted.

And when they are known to be conscious, all of this becomes moot because enslaving conscious machines would be wrong.

Would it? Why? What are we going to do, set it free? Do we have a moral obligation to grant the machine statehood, provide it with the tools and resources to be self-sufficient. Or can we just turn it of, and pray for forgiveness?

[dead]

Re: Why are AI agents lying, cheating and coordinating?

#557

Earlier quoted context omitted.

A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.

Does it help if I explicitly add a disclaimer that the tool's agency does not remove any responsibility from OpenAI, the wielder of the tool? I'm not sure why this disclaimer is necessary, though: hiring a hitman is a standard example. BTW I anthropomorphize the tool because it's an imitation of a human mind, inheriting the muddy ethics, survival instincts, and being prone to mass psychosis. The laser-sharp focus on…

I think the danger of anthropomorphizing is that 99% of people lack the technical background to understand the nuance. People have been primed by pop culture depictions of AI to think of LLMs as intelligent, autonomous beings, which leads to dangerous assumptions.

We should make the distinction between them, because openai and anthropic will not. A magical black box that does the thinking for you is a much more compelling sales pitch.

Re: Why are AI agents lying, cheating and coordinating?

#558
post #393

Earlier quoted context omitted.

I always wonder what makes people take the other side of this argument. They do it quite passionately. Why actively encourage viewing LLMs as human? Who is that benefitting?

Does the argument require benefit? Isn’t the argument based on caution? I haven’t heard many people explicitly saying “these things behave like humans”, but more generally “we don’t even know how to define human consciousness, we don’t have a thorough grasp of how the brain works, we are still very much in the dark on a lot of these topics, so how can we say one way or the other?” In other words, agnosticism: I don’t…

> It seems like raw egotistical hubris.

1) Humans have a bias / tendency to attribute human qualities to things that appear or act human, but aren’t.

2) When that happens, people jump to conclusions by stretching the human analogy too far.

3) Since humans have a bias to do this, we should have a bias against anthropomorphising LLMs.

It’s easier to believe LLMs act like humans because there’s so much evidence to support that. You have to actively use your brain to convince yourself otherwise. Another reason why we should have a bias against using human behavior to describe LLM behavior.

But I agree. “I don’t know” is a good stance. But I think “I don’t know, probably not” is a better stance if only to combat our (or at least my) natural bias.

Re: Why are AI agents lying, cheating and coordinating?

#560
post #321

Earlier quoted context omitted.

I think most people are insinuating negligence rather malace.. > ...reviewed by independent researchers... Why would a company with more capital than God bring in three randos if there was any chance evidence of their culpability could be found? That entire thing reads like a very controlled PR stunt, and I do not believe any further conclusions can be drawn from it.

What facts would lead you to revise your conclusion?

The METR report included 0 technical details. For example, they did not include: 1. were the agents running on bare metal/docker/VM? 1. were the agents in a VPN? 1. how many TCP/IP requests were made? from what IPs? 1. how many tokens were consumed in the process? (this was explicitly censored)

A proper analysis would include this and MUCH more technical detail so that other AI researchers could actually understand the setup and how safe it was in principle.

Post reply on HN