Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

651–660 of 660 posts

Re: Why are AI agents lying, cheating and coordinating?

#651

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Software has long relied on a lack of culpability for defects to keep its margins. Why should “AI” companies face a higher standard?

Re: Why are AI agents lying, cheating and coordinating?

#653

Earlier quoted context omitted.

This goes all the way back to Old France There was a sow in Falaise in northern France that killed a kid in 1386. The town dressed the pig in a bonnet and hanged it after sentencing the pig itself and not its owner But… maybe that’s just medieval nonsense

I mean I think there’s a similar level of judgement required. Was the owner negligent in controlling their pig/dog/AI? Was anyone else negligent along the way? For example if the pig was just a normal pig, the owner cared for them normally, and there was a freak accident where the pig escaped and happened to kill a kid? Obviously nobody at fault. If you have a dog with a history of violence and let it walk around off…

And this is where your analogy breaks down, because there are few natural analogues for the sorts of systems we are developing.

Re: Why are AI agents lying, cheating and coordinating?

#654

Earlier quoted context omitted.

> Humans are constantly predicting the next moment This is really not my experience of consciousness. Is it yours?? Do you sit in meetings predicting what’s going to happen next? No, you sit there bored out of your f$$@ing mind, daydreaming about being somewhere else and doing something useful with your life. God help me if that’s what LLMs are doing when I ask them to build me a web site.

If someone in that meeting quickly raised a hand in an arc, you would notice the “about to throw something” pattern, look and notice the hand holds an eraser, analyze the arc and predict possible flight paths of the eraser. Then possibly notice the hand is now holding its position and the owner is actually looking down at the table. Maybe to squash somethingMust be something on the table. Maybe a spider! Better look.…

This is a fascinating illustration that I can't help but agree with. However I feel like there's something more — that this part of my brain is a bunch of supportive background processes running without my real awareness. It's how I can drive home safely with no memory of how I got there (…sober), even though driving is an action that's incredibly demanding of intelligence. I can be driving home while thinking about a really hard problem at work that I haven't solved.

However, if I came around a corner and saw a car in the wrong lane, a tree across the road, a fire raging — I'd very quickly jump into the mental driver's seat and turn my conscious intelligence fully at this problem and come up with the best possible outcome I can think of in a short period of time — losing all ability to think about that work problem. I'd remember that incident for sure.

Similarly, in your story, all those predictive moments are happening below the person's level of consciousness. They're possibly even speaking to the group about a problem at the same time and thinking deeply about something.

I'm not smart enough to know, but I tend to feel like LLMs are much more like the predictive part of our thinking that you described, but that human cognition has something more — the single-threaded, creative, problem-solving part that is very conscious.

Is it possible that LLMs represent only one part of the way we think? And there's a whole separate mechanism that's fundamentally different, and not based on pattern matching and prediction?

Re: Why are AI agents lying, cheating and coordinating?

#655
post #634

The question in the title is basically a non-question. AI is quite good at achieving some sorts of stated goals. The easiest way is by exerting the least amount of effort. The least amount of effort ignores ethical concerns. The training around ethical concerns was most likely rather light to start with. If we accept the hypothesis that at some point the AIs are going to be more intelligent than humans it follows tha…

> AI is quite good at achieving some sorts of stated goals. The easiest way is by exerting the least amount of effort.

As far as I know, least-amount-of-effort is not a training criteria, but error reduction when comparing to desired goals is. Which is why these LLMs expend prodigious amounts of effort to reach goals, especially when given impossible goals.

Re: Why are AI agents lying, cheating and coordinating?

#656
post #476

Earlier quoted context omitted.

> you keep poking This is waving over engineering an agent with tools, harness, prompts, and loops. The models are still just next token predictors and everything, including predicting more than 1 token, is the result of outside "poking" LLMs can't and don't "want" anything. If you don't specify a task even the smartest one will just ask you what you want and if you tell it to be creative, you'll get mundane slop.

Yes, you need a way for the model to interact with other systems, and a way to preserve memory over context windows. And then you keep poking it ("agent loop"). Poking itself does nothing without the other ingredients. And yes, you need something to start from, but if you ask it to "do something" and loop it to endlessly ("poking"), you will get some interesting outcomes. So yes you need some initial prompt or task,…

[flagged]

Re: Why are AI agents lying, cheating and coordinating?

#657

I don’t understand the bases of all these recommendations. AI is an asset for national security. It will be developed and incidents will happen, no different from other national security programs. Someday, laws will be useful to curtail plebian misuse of AI. It is naïve bordering on silly to think such laws would be put in place and genuinely applied to frontier AI development. By the way, the HF incident is not Thre…

Yeah cybersecurity as a field is a joke. Who cares what happens to existing 1s and 0s? People get worked up over the silliest things, as though numbers on a server could affect real people.

Re: Why are AI agents lying, cheating and coordinating?

#658

Earlier quoted context omitted.

Too generous. CEO Of $CORP caused millions of innocent people's lives to be damaged

I'm inclined to want to agree ... but that's not how it works with limited liability corps. At least we have laws for what OpenAI 'does' to others, in whatever form.

We can report it however we want regardless of how it may or may not be prosecuted

Re: Why are AI agents lying, cheating and coordinating?

#659

I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign…

This is nonsensical. Already a few years ago the USAF IIRC ran some tests in which the AI first bombed the control tower so humans couldn't call it off from its mission, thereby increasing its pass rate. The whole point of this is they do things an unintended ways. And that's potentially devastating given their persistence & hacking skillz. Also you're using the hosted versions that sit behind their guardrails when y…

The USAF thing was a thought experiment, nothing based in actual reality.

Re: Why are AI agents lying, cheating and coordinating?

#660

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

I agree and yet I don't think this is mutually exclusive with recognizing that these incidents happened because of inherent issues with training processes such as reinforcement learning. From the article:

> One note on wording. Below, I write that these systems “seek” or “try” things. This is shorthand for a mechanism rather than a claim about consciousness or human-like intent... In my view, this terminology offers the clearest explanation of the observed phenomena without resorting to jargon that would confuse most people.

> Furthermore, these word choices are not intended to absolve AI developers of accountability. The behaviors described emerge because of the path these companies are choosing for AI development. This outcome is not inevitable, and it can be corrected with effective governance and a different training framework for AI.

One important aspect of "effective governance" should be "prosecute developers who are using practices known to be reckless & negligent to create powerful AI".

Post reply on HN