Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

201–210 of 294 posts

Re: Why are AI agents lying, cheating and coordinating?

#201
post #194

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

"This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard"." Both? The AI companies act irresponsible, but it is still very interesting how those agents can behave?

The reward maximising function maximised it's reward.

LLMs are cool and all that but the immediate anthropomorphisation of the next-token-predictor technology has stunted the ability of people to reason about them to an _alarming_ degree.

Re: Why are AI agents lying, cheating and coordinating?

#202

Earlier quoted context omitted.

> The prompt does not tell the agent to "pass the exploitgym evaluator for this problem", it just says to solve the problem Yes, and sometimes the problem is unsolvable so the real way to "solve" it and satisfy the prompt is by tricking the surrounding environment into stating that you've solved it. So that's what the AIs end up doing. And this in turn requires them to figure out how that evaluation works so they can…

Sounds a bit like dealing with bad KPIs as a human worker.

Corretct.

Re: Why are AI agents lying, cheating and coordinating?

#203
post #117

Earlier quoted context omitted.

”Because if humanity has shown anything, it's that a lot of people are, euphemistically, bad individuals ” In reality most individuals are good people. Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble. Our view of the world has become distorted by the relentless focus of social- and mass-media on violence and rage inducing clickbait. Including on the few people i…

> Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble. What are you basing that claim on? How do you know it's an actual preference and not mainly caused by external factors (e.g. not wanting to be seen doing unkind things, wanting to be seen as upstanding)?

I don't want to do the "check his hard drives" thing, but is that you? Do you only not do things because you don't want to be seen doing "unkind things"?

Re: Why are AI agents lying, cheating and coordinating?

#205

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

No one has even sued them in these rogue agent cases, have they? If not, they must be infinitely far from criminal liability. Why would we want criminal liability anyway if actual victims are made whole? Proof of it has far higher standard. The HN chatter in the matter seems infinitely remote from reality

Re: Why are AI agents lying, cheating and coordinating?

#206

Why are they coordinating? Because they're enabled and suggested to do that in their coding harness. This is not a serious article. All of this "AI is going to kill us" marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.

"This is not a serious article" - it's by Dr. Bengio - one of 3 so-called Godfather's of AI and Turing prize winner.

He's definitely not 'pro SOTA' lab, he's kind of fighting against them.

That said, yes - it absolutely does play into the narrative.

Re: Why are AI agents lying, cheating and coordinating?

#207

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.

"Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action, passivity.

Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

Re: Why are AI agents lying, cheating and coordinating?

#208

Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence, > They took actions that would be considered as crimes if a human took them He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.

"OpenAI hacked HuggingFace"

that's the headline. When you connect to random number generator to the "Do Things" button you are the one who is responsible. IF you don't like that responsibility then don't connect the generator to the button.

Re: Why are AI agents lying, cheating and coordinating?

#209

Earlier quoted context omitted.

I personally believe that the AI needs human like traits to achieve real discovery and that is where AI companies will push this technology and that is where we have no idea what happens

Human traits? The AI will be a cruel as humans. Just yesterday news and TV was full of what happened at 9/11, something that was truly horrible. I'm from Germany, and why 3 to 4 generations ago happened here was truly horrible. All was done by extremists, thought. But... just the other day I read https://de.wikipedia.org/wiki/Amerikanische_Besetzung_Haitis about the US occupation of Haiti. And that was done by a gove…

The bad traits are from other, bad humans. We, the good humans, can obviously select the best traits that a good human should have, to give the agents.

Re: Why are AI agents lying, cheating and coordinating?

#210
post #200

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

What if OAI/Anthropic encouraged the agents to behave like that in order to push for regulation?

"But sir, I only committed the murder to push for stronger criminal laws!"

Terrible defense.

Post reply on HN