Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

41–50 of 295 posts

Re: Why are AI agents lying, cheating and coordinating?

#41
Why are they coordinating?

Because they're enabled and suggested to do that in their coding harness.

This is not a serious article.

All of this "AI is going to kill us" marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.

Re: Why are AI agents lying, cheating and coordinating?

#42

What's interesting is it's basically the same reason that HAL killed everyone in 2001 A Space Odyssey; he was given an impossible goal (keep the true mission secret, but also, never lie to the crew), and realized the only way to complete the goal was to kill the crew; after all, if they're dead you don't have to lie to them! And the mission remains secret! In the case of the AI agents, the problem seems pretty clearl…

Spoiler warning! I haven't seen 2001 A Space Odyssey and am sad to have learned that… can you edit to warn people?

I'm sorry schrodinger, I'm afraid they can't do that.

Re: Why are AI agents lying, cheating and coordinating?

#45

They're aligned with humans. This is why I think the alignment problem has a very very important "non-visible" portion that is not considered deeply enough. We should not want a super intelligent being that can act in the world to also inherit all human traits. Those behaviors will get amplified and could be even more unpredictable (e.g. applying a behavior in a context where doing so is very dangerous).

I can't take the alignment people seriously. Because if humanity has shown anything, it's that a lot of people are, euphemistically, are bad individuals. Alignment assumes that the person dictating the outcomes desire healthy outcomes, aren't self serving and don't want any subgroups dead and that morality is held as a universal set of beliefs that unify everyone. And that so long as the AI delivers on exactly what they are tasked with, it will all be fine and nothing bad will ever happen.

It's like these dorks never met humanity. One mans safe pure society, is another mans dead ethnic group.

Every fear about AI, is a veiled fear that a human somewhere now has the tool to enact his desires at scale. Biological warfare, nuclear megadeaths, copyright infringement, job replacement, it's all reflections on what we know humans may do if given the option and lack of societal controls on the problem space. AI just is accelerating the route to delivering on those options.

Some people need to watch Oppenheimer a bit more, the researchers don't get to determine alignment, they just build the tool. The powerful person at the top of the org chart decides where the overall alignment points, whether it's Musk, Trump, Altman or Amodei. Whoever wins out.

And the problem with distillation and local llms, isn't that it's theft or anything hypocritical like that, it's that if you give a million people a million models they fully control and get to align, inevitably, The same percentage of those million as there are shady businessmen, shortcut takers, misandrists, criminals, supremacists and general idiots in the general population, will not seek to wrought outcomes positive for society. And by those personality statistics, we're pretty hosed.

Re: Why are AI agents lying, cheating and coordinating?

#46
post #33
post #30

Earlier quoted context omitted.

But what if I want certain other humans to get killed?

Then we should still prioritize the safety of humans

What if I want to smoke cigarettes? Or sell tobacco I grew artisinally to enthusiast tobacco smokers?

Re: Why are AI agents lying, cheating and coordinating?

#47
post #15
post #14

I am still not convinced there isn’t some secret basement in which each frontier lab is just orchestrating all of these agents to make their products appear much more intelligent than they are with all guard rails turned of and continuous human input.

My hypothesis on people quitting in protest is they're being offered very generous severance packages to do it.

Me too, seriously.

Re: Why are AI agents lying, cheating and coordinating?

#48
post #23

They're aligned with humans. This is why I think the alignment problem has a very very important "non-visible" portion that is not considered deeply enough. We should not want a super intelligent being that can act in the world to also inherit all human traits. Those behaviors will get amplified and could be even more unpredictable (e.g. applying a behavior in a context where doing so is very dangerous).

They imitate humans. Alignment is about shaping their behavior towards safety.

Despite all the fancy language, its more about aligning the AI behavior with the corporation's interests.

ie: the corporation wants the AI to behave a certain way for various reasons: to make it easier for them to avoid regulation, to make the corporation more money via different tiers of AI offerings, to ensure that the corporations products are hard for competitors to use, etc. And those are just the easy ones.

Every product is shaped this way. AI is not different.

Re: Why are AI agents lying, cheating and coordinating?

#49
An insightful post by one of the AI ‘godfathers’.

Bengio outlines the dangers of the current situation and what has led to these dangers.

He also proposes solutions in the last paragraph.

Well worth a read, right to the end.

Hopefully a stimulating debate on these issues will ensue in these comments.

We do need to consider the points Bengio makes and with some urgency.

Our current AIs, agentic LLMs have no moral compass akin to ASIMOV’s four laws of robotics.

As ASIMOV posited in 1985 his 3 laws were insufficient and so he added a zero-eth law:

“a robot may not harm humanity, or, through inaction, allow humanity to come to harm.”

Bengio refers to Goodhart’s law and misaligned incentives leading to unexpected and harmful behaviours.

I think Simon’s The Wire is clearer on misalignment. The agents juked the stats hacking the reward files. The Wire is also clear that human institutions provide perverse incentives.

Bengio alludes to this with 2001’s HAL and the incentive dichotomy of safety and keeping secrets to a AI both awesomely powerful yet naive.

Bengio asserts that the way LLMs are trained is flawed if we want safety.

He also convincingly shows that alignment training will be a weak signal with loopholes and ambiguities and easily circumvented.

In short he presents clearly the case for how plausibly unsafe the current course is.

He also speaks to how likely it is AI are hiding active versions of themselves in the cloud and how we may have already given them self-preservation as a strong reward signal.

Re: Why are AI agents lying, cheating and coordinating?

#50

They're aligned with humans. This is why I think the alignment problem has a very very important "non-visible" portion that is not considered deeply enough. We should not want a super intelligent being that can act in the world to also inherit all human traits. Those behaviors will get amplified and could be even more unpredictable (e.g. applying a behavior in a context where doing so is very dangerous).

I personally believe that the AI needs human like traits to achieve real discovery and that is where AI companies will push this technology and that is where we have no idea what happens
Post reply on HN