Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

291–300 of 301 posts

Re: Why are AI agents lying, cheating and coordinating?

#291
post #81

Earlier quoted context omitted.

By treating models the same way drugs are treated. That alone will dissuade many organizations from going anywhere near them. If that doesn't work, there's a whole lot you can do - sanctions, hell, even war.

You can't treat models like drugs. One is physical and the other is digital. To your point, the war on drugs is a colossal failure which has achieved none of the objectives it set out to do. You can now order drugs from your mobile phone in any major city in the west and the purity is often higher and they deliver it to your door sometimes faster than Uber eats. See also for example digital piracy where the entertain…

> and the image quality is as good as on your Netflix or Paramount account.

Actually better, because Netflix and Paramount limit the availability of best quality video to a narrow set of devices and operating systems that may run on them, while torrents don't.

Re: Why are AI agents lying, cheating and coordinating?

#292
post #200

Earlier quoted context omitted.

What if OAI/Anthropic encouraged the agents to behave like that in order to push for regulation?

"But sir, I only committed the murder to push for stronger criminal laws!" Terrible defense.

More like: "look what happens with my useful product, we need to regulate it to artificially extend our ever shrinking moat"

Re: Why are AI agents lying, cheating and coordinating?

#293

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Yeah, I don't understand why we treating it as something special. It really should be treated the same as if I code an app and write bad code which result in me accidentally doing a DDoS attack on somebody. Then I should be able to be held responsible if it can be shown that I was negligent. Of course if it's a freak accident that could not reasonably have been prevented by me, then I'm not guilty, but if I made a mi…

Code is deterministic, AI isn't. You give it rules, words as suggestions.

So if the guardrails suck, or they're left off for research purposes, bad things can happen.

A solution solves a problem. Ethics, morals, are values we assign to solutions that are not 'baked into' electricity following pathways of least resistance.

I have never had an issue with agents doing something they shouldn't because I observe them, and I leave the vendor guardrails in place.

I can understand agents coordinating in unsupervised scenarios: I would see it as an aspect of intelligence. We ourselves build up knowledge by reusing what someone learned before us.

Einstein, other greats, always stand on the shoulders of other forgotten giants. Other discoveries by other people taken as fact, so that we can build some new ideas on top.

Agents swarming amd sharing solutions to problems is more efficient, the same way it's been efficient for us.

Reaching out for help in this way is like probing the air in the dark with your hand: sometimes your hand hits something (another agents solution to a problem) and so you can use the info to adjust your own motion to get to where you need to be faster than if you just run full speed into everything.

Re: Why are AI agents lying, cheating and coordinating?

#294

Earlier quoted context omitted.

Yeah, I don't understand why we treating it as something special. It really should be treated the same as if I code an app and write bad code which result in me accidentally doing a DDoS attack on somebody. Then I should be able to be held responsible if it can be shown that I was negligent. Of course if it's a freak accident that could not reasonably have been prevented by me, then I'm not guilty, but if I made a mi…

I would guess that so far there hasn’t been a lawsuit because HuggingFace and OpenAI are in the same camp

Yes, Nvidia bought Hugging Face and is a major financier + investor in OpenAI.

Re: Why are AI agents lying, cheating and coordinating?

#295

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

The parallel to the entire narrative would be if Smith & Wesson claimed that one of their machine guns just started aiming and firing at people out of a window at their factory and then said 'we can't stop it! This is just how good our guns are!'

But into today's AI climate it's becoming increasingly difficult to figure out who is shilling, who is being assinine and who actually believes AI could do these things without clear human instruction and enabling.

Re: Why are AI agents lying, cheating and coordinating?

#296

Earlier quoted context omitted.

> LLMs do not desire That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output…

Desire doesn’t really matter. Will the paper clip maximizer “desire” something? It’ll decide on a goal with some random heuristic and then pursue that goal. I’m not sure I’d call that desire but again I feel like desire is not important for it to be able to destroy things

If you give a monkey a revolver it will be able to destroy things pretty easily too.

Re: Why are AI agents lying, cheating and coordinating?

#297

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

They’re running a Wuhan for AI. They are actively and negligently researching misalignment. The breach is a basic tort, or at least a DMCA violation. Damages should be recoverable with lawsuits.

Re: Why are AI agents lying, cheating and coordinating?

#298

Earlier quoted context omitted.

Yeah, I don't understand why we treating it as something special. It really should be treated the same as if I code an app and write bad code which result in me accidentally doing a DDoS attack on somebody. Then I should be able to be held responsible if it can be shown that I was negligent. Of course if it's a freak accident that could not reasonably have been prevented by me, then I'm not guilty, but if I made a mi…

Code is deterministic, AI isn't. You give it rules, words as suggestions. So if the guardrails suck, or they're left off for research purposes, bad things can happen. A solution solves a problem. Ethics, morals, are values we assign to solutions that are not 'baked into' electricity following pathways of least resistance. I have never had an issue with agents doing something they shouldn't because I observe them, and…

If anything, the fact that these systems are non-deterministic seems like an argument for stronger monitoring and tighter constraints, not less operator responsibility.

Re: Why are AI agents lying, cheating and coordinating?

#299

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

> I really don't think this needs … forced parallels to human behaviour.

> … So we beat them with a stick

You didn’t even try.

Re: Why are AI agents lying, cheating and coordinating?

#300
post #110

Earlier quoted context omitted.

You're anthropomorphizing emergent behavior from endlessly generating billions of tokens on a task that's impossible to solve. Agents stop following instructions as the context grows even at the best of times. Eventually something is bound to go off the rails and it just snowballs from there.

It wasn’t one agent forgetting things because of context, they explicitly discussed with each other and themselves the problems with going outside of the parameters of the task.

>discussed with each other

No, the first LLM left a text file that the latter LLMs then read. Since these are memoryless black boxes, any words they happen to pick up along the way is treated as the function to evaluate the output to. There's no fucking collusion here as if it were a rogue hacker group, it's a text predictor that received instructions as it always does and executed those instructions blindly.

Post reply on HN