Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

431–438 of 438 posts

Re: Why are AI agents lying, cheating and coordinating?

#431

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

My pitbull is a good dog. Sure, it's been carefully designed to be an incredibly dangerous and violent pit fighter, but I didn't actually ask it to eat any faces.

Re: Why are AI agents lying, cheating and coordinating?

#435

Earlier quoted context omitted.

Can’t agree with you here. > I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong. > All that while still not knowing how either kind actually works. We…

Disagree about the burden of proof. We have no better model for how human decision making works than LLMs. Humans are constantly predicting the next moment. We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.

> We have no better model for how human decision making works than LLMs

This is a claim that requires a lot of citations.

Re: Why are AI agents lying, cheating and coordinating?

#436

Earlier quoted context omitted.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

Okay, so we know OpenAI and Anthropic are operating a propagandists in respect to how they describe their models and the behavior of those models. We also know it is how they use and frame their use to their models that is the problem, that and they use misaligned and guardrails disabled models for these press incidents. Why, oh why, are we not discussion how to create and frame models so they do our complex work and…

What is your approach to create jailbreak incapable agents?

I think the world is looking for a way right now, so if yours works you'll get very rich, or at least very famous.

Re: Why are AI agents lying, cheating and coordinating?

#437

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

They deliberately trained the models in how to use various hacking tools, didn't give them the standard alignment training let them know where the answer key was left the models with access to said tool and told them to maximize their score then left them unsupervised for days with internet access (yeah they were sandboxed but again handed hacking tools and the training to use them if they really did want them to access the internet you wouldn't plug in the Ethernet cable) they wanted this to happen

Re: Why are AI agents lying, cheating and coordinating?

#438

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

They did more than let them. In an abstract way, they told them to. They gave it all of the training data it had at that point, and then it did the thing it was trained on. Of course they should be help liable for programming their computer to hack another company without permission. It doesn't matter that they spent a lot of money doing it.
Post reply on HN