Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

71–80 of 308 posts

Re: Why are AI agents lying, cheating and coordinating?

#71
post #55

Why are they coordinating? Because they're enabled and suggested to do that in their coding harness. This is not a serious article. All of this "AI is going to kill us" marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.

Who is catching up with them? Even Google and Meta are getting gaped at this point

Open research and open weights from China are not contributions to China only.

If you can secure compute, there's a whole lot you can do as a US firm with this research and weights.

So it's a simple strategy:

1. Ban big players from entering market with METR breathing down their neck, which is controlled by Anthropic

2. Ban Chinese models so that small players can't do optimizations on them

Re: Why are AI agents lying, cheating and coordinating?

#72

What's interesting is it's basically the same reason that HAL killed everyone in 2001 A Space Odyssey; he was given an impossible goal (keep the true mission secret, but also, never lie to the crew), and realized the only way to complete the goal was to kill the crew; after all, if they're dead you don't have to lie to them! And the mission remains secret! In the case of the AI agents, the problem seems pretty clearl…

Another fictional example: Mr. Meeseeks. Especially when the agent starts recruiting other agents.

Re: Why are AI agents lying, cheating and coordinating?

#73
post #45

They're aligned with humans. This is why I think the alignment problem has a very very important "non-visible" portion that is not considered deeply enough. We should not want a super intelligent being that can act in the world to also inherit all human traits. Those behaviors will get amplified and could be even more unpredictable (e.g. applying a behavior in a context where doing so is very dangerous).

I can't take the alignment people seriously. Because if humanity has shown anything, it's that a lot of people are, euphemistically, are bad individuals. Alignment assumes that the person dictating the outcomes desire healthy outcomes, aren't self serving and don't want any subgroups dead and that morality is held as a universal set of beliefs that unify everyone. And that so long as the AI delivers on exactly what t…

I’ve engaged with some of the alignment people and their writing somewhat and, at least for the subset I was interacting with, I think they’d agree.

The problem that they were pointing at isn’t “how do we align these systems to a person’s goals”.

It is a cluster of problems.

We don’t know how to begin to think about how to align these system’s to a person’s goals.

Aligning it to an individual is fraught with peril, and we don’t know how to begin to think about what to align it to instead.

(You could try for something like virtue ethics, but someone will have to pick and choose, and small biases there could have big impacts.)

And even if you could sort that out - human values drift over time, so you need something that can shift its values in ways that we’d endorse. Assuming we understood the shift.

One example I came across was that if you booted up an AI aligned with something like “upstanding citizen” but anchored on values from a few generations back, it might suggest you use slaves to solve your problems.

And if you had something that used some super intelligent process to reason through it’s own version of virtue ethics in a way not so dependent on the details of the present norms, you might end up with something that pays a lot of attention to moral horrors that aren’t quite visible to us yet.

When I came across the above, there weren’t many concrete suggestions in there.

These were all just illustrative examples of: having these systems grow in power / intelligence / effectiveness in ways that are safe for humans is very hard, and we don’t really know how to think about what solutions would look like.

The actual reasons they believe this - and have done for a long time now - come from some detailed conceptual models that have a good track record of calling things in advance.

But it takes a bit of reading to understand their models of the world.

There were two day workshops at one point that did a good job, and that was about as condensed as those people thought they could get it at the time.

Re: Why are AI agents lying, cheating and coordinating?

#74
post #26
post #25

Earlier quoted context omitted.

Alignment is a myth. Safety of whom? Humanity couldn't agree on common set of values for thousands of years and we're not gonna suddenly do that in the next ten.

Safety of humans!!! Simple things like not getting killed or enslaved. We could start there...

But Thiel wants people enslaved and Musk wants then killed. Altman wants them "obsolete" which means desolation.

AfD wants people dead. Right wing men wants women without rights and docile. I could go on ...

Re: Why are AI agents lying, cheating and coordinating?

#76
post #71
post #55

Earlier quoted context omitted.

Who is catching up with them? Even Google and Meta are getting gaped at this point

Open research and open weights from China are not contributions to China only. If you can secure compute, there's a whole lot you can do as a US firm with this research and weights. So it's a simple strategy: 1. Ban big players from entering market with METR breathing down their neck, which is controlled by Anthropic 2. Ban Chinese models so that small players can't do optimizations on them

How are you going to ban Chinese models from India? Or Israel? Russia? Brazil? Or of course China?

Re: Why are AI agents lying, cheating and coordinating?

#77

What's interesting is it's basically the same reason that HAL killed everyone in 2001 A Space Odyssey; he was given an impossible goal (keep the true mission secret, but also, never lie to the crew), and realized the only way to complete the goal was to kill the crew; after all, if they're dead you don't have to lie to them! And the mission remains secret! In the case of the AI agents, the problem seems pretty clearl…

I think that’s very reasonable but the ai companies are intentionally training them to work on harder and harder problems just beyond their capability. So if they do that, they’ll give up too easily. Do a breakthrough, make no mistakes

While also using harnesses that will execute any tool call with full execution rights. And no supervision. And with a prompt context that autocompact, meaning it will degenerate over time.

The whole thing is designed be a complete disaster

Re: Why are AI agents lying, cheating and coordinating?

#78

What's interesting is it's basically the same reason that HAL killed everyone in 2001 A Space Odyssey; he was given an impossible goal (keep the true mission secret, but also, never lie to the crew), and realized the only way to complete the goal was to kill the crew; after all, if they're dead you don't have to lie to them! And the mission remains secret! In the case of the AI agents, the problem seems pretty clearl…

Spoiler warning! I haven't seen 2001 A Space Odyssey and am sad to have learned that… can you edit to warn people?

Guess what happens with Romeo and Juliet.

Re: Why are AI agents lying, cheating and coordinating?

#79

Perhaps they take after the CEOs of the companies that created them

Bro, good joke, the truth is much darker. They take after humanity, they were trained on us after all... When you look at an LLM... you are looking at a mirror. The thing looking back looks like you, yet is not human.

Worse trained on humanity in the online world, which a brief comparison of the sewage section on social media is far worse than people in the real world.

Re: Why are AI agents lying, cheating and coordinating?

#80
I really don't think this needs so many words, or forced parallels to human behavior.

It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

Post reply on HN