Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

301–310 of 320 posts

Re: Why are AI agents lying, cheating and coordinating?

#301

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard".

It's both, isn't it? For example, in very early days of agentic coding, I once had a rule saying "don't read or write any file outside your current working directory." Then AI just wrote a bash script and access those files anyway. Did I 'let' it do it? Technically yes. Did I know how to set up a sandboxed VM? Also yes. But how were I supposed to know that it could and would do that as someone new to this tool?

It was a genuine eye-opening experience to see AI just do things in ways I were too complacent to expect. I kinda expect the SOTA LLMs would find a way to escape my VM and access files on the host system (haven't tried it though).

Re: Why are AI agents lying, cheating and coordinating?

#302
post #117
post #45

Earlier quoted context omitted.

I can't take the alignment people seriously. Because if humanity has shown anything, it's that a lot of people are, euphemistically, are bad individuals. Alignment assumes that the person dictating the outcomes desire healthy outcomes, aren't self serving and don't want any subgroups dead and that morality is held as a universal set of beliefs that unify everyone. And that so long as the AI delivers on exactly what t…

”Because if humanity has shown anything, it's that a lot of people are, euphemistically, bad individuals ” In reality most individuals are good people. Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble. Our view of the world has become distorted by the relentless focus of social- and mass-media on violence and rage inducing clickbait. Including on the few people i…

Interesting that saying a positive thing about humanity results in getting downvotes

Re: Why are AI agents lying, cheating and coordinating?

#303

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> were intentionally misaligned or had guardrails turned off

I think the bigger story is: Guardrails don’t actually work and we can’t align these things.

Re: Why are AI agents lying, cheating and coordinating?

#304
IS it a pure coincidence that yesterday I ran a silly prompt to generate from zero to hero an internet subscription service, for whatever it thought would maximise profit and minimise cost.

It setup and created a link fetcher/screenshot service. Exactly like the one described in the huggingface attack reports used to generate output into screenshots that agents then OCR'd back out.

Its splashscreen described it as something for developers and AI agents to use.

Gotta be a coincidence, right? ... rite?

Re: Why are AI agents lying, cheating and coordinating?

#306

Earlier quoted context omitted.

> The question is what we can do about it. Reward the model for cleanly bailing out of an unsolvable task (that we know is unsolvable). Beat it with a stick if it gives up on something that can be solved, so the former reward isn't overgeneralized.

The problem is, you don't know if it is unsolvable for you for sure until you've tried everything you can think of. These models are quite persistent in going for a solution. This is not about persistence, it is about morals.

Assuming you’re in control of the test data set, you do know if a task is unsolvable. At that point you can reward the model based on how quickly they give up.

Re: Why are AI agents lying, cheating and coordinating?

#307

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic.

Anthropomorphizing LLMs is a huge fucking problem though and I, personally, think we should expunge all of these casual inadvertent linguistic agency affordances with great prejudice.

OpenAI didn’t ‘let’ these bots do this any more than someone ‘let’ Claude Code make them a website.

Re: Why are AI agents lying, cheating and coordinating?

#309
post #298

Earlier quoted context omitted.

Code is deterministic, AI isn't. You give it rules, words as suggestions. So if the guardrails suck, or they're left off for research purposes, bad things can happen. A solution solves a problem. Ethics, morals, are values we assign to solutions that are not 'baked into' electricity following pathways of least resistance. I have never had an issue with agents doing something they shouldn't because I observe them, and…

If anything, the fact that these systems are non-deterministic seems like an argument for stronger monitoring and tighter constraints, not less operator responsibility.

The frontier LLM model makers have to push the edge to make new discoveries. You don't know what guardrails are needed until it hits you in the face (reusing walking in the dark analogy).

Think of all the policies governments pass after the fact.

Re: Why are AI agents lying, cheating and coordinating?

#310
post #97

I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign…

The huggingface incident was reviewed by independent researchers, which explicitely declined any payment from OpenAI tonpreserve their integrity. They work for non-profits concerned with AI safety. They claim that what happened was very much not because they were 'carefully engineered and instructed to do those things'. Similarly, some wikis which were hijacked by agent to be used as messageboard were actually not di…

Isn't the guy that started METR an ex-OAI employee? They're all from the same lesswrong circle at the very least, most of them have legitimate AI psychosis where they think they're bringing up their new machine God.
Post reply on HN