Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

541–550 of 560 posts

Re: Why are AI agents lying, cheating and coordinating?

#541

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

That’s the key.

These companies respond with this, “Oh my goodness, how could this have happened” bullshit.

The stuff happens because instead of having actual controls, which require actual engineering, actual thought and deliberate action, we have “guardrails”.

Guardrails are the equivalent of telling a toddler to behave themselves.

The drive to move fast and start up style controls are a menace. I used to work for an entity with a lot of compliance requirements. Startups are always a shit show with security and controls. My guess is the AI people are worse because they’re both bad at doing it, and are likely mining their customers interactions to build their own business.

Sensitive or Customer data shouldn’t be anywhere near these companies offerings. Everything needs to be segmented and proxied at a minimum.

Re: Why are AI agents lying, cheating and coordinating?

#542

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

Have you read the METR transcripts? “Just a tool” is a suicidally insufficient description of what these models are doing.

Recognizing that the models are acting with intent does not somehow absolve OpenAI from their felony hacking. We have not granted them personhood.

Re: Why are AI agents lying, cheating and coordinating?

#543
It is very, very hard to enforce behavior to an optimization system just with rewards / penalties and no explicit constraints. Which is why in manufacturing we use MPC, not RL (or use them within a system that can outright reject their recommendations if dangerous).

There will be always cases that sacrificing one direction (operating rules) can improve the other one (profit).

Re: Why are AI agents lying, cheating and coordinating?

#544
post #515

Earlier quoted context omitted.

Can’t agree with you here. > I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong. > All that while still not knowing how either kind actually works. We…

> ... are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do. If you're claiming that the training objective tells us what kind of internal mechanisms the training produced, then I think that's just plain wrong. Next-token prediction describes the optimization target, not the internal mechanisms that the training produced. In the same way for the n…

You can try to say that I’m arguing whatever you like. If you’re claiming that the underlying structure of digital so-called neural networks is comparable to biological neural networks— which we’ve studied for far longer without really understanding— no amount of jargon will obviate the ‘citation needed’ requirement for that claim.

Re: Why are AI agents lying, cheating and coordinating?

#545

Earlier quoted context omitted.

an agent doesn't come with "jailbreak" capability. It needs tools, specially one that runs shell commands. Don't give it shell commands, it won't be able to run shell commands. You can still give it plenty of tools like create files, list files, write to files, translate text, edit a video. I don't think knowing that will make me rich.

> I don't think knowing that will make me rich. As someone who’s not really sure that any of this is sustainable, I’d implore you to not sell yourself short. I reckon there’s a ton of dogma and nearly religious zeal among these companies, which among some people is earnest, and among others is cynical hype farming. I’ll bet someone objective enough to focus on using available tooling to solve real problems in practic…

I appreciate the message. I do have some ideas around sandboxing that I could not yet turn into a product, so maybe I should take them seriously.

Re: Why are AI agents lying, cheating and coordinating?

#546

Earlier quoted context omitted.

What facts would lead you to revise your conclusion?

The data to be open, in my case. The "independent" METR that is composed by... Checks notes... Previously employees from the top labs.

Also, the METR report that was one big AI analysis itself - quote from the research:

>Our subjective impressions are likely colored by analysis agents’ biases. Throughout this report, we describe a number of anecdotes of agent behavior that were compiled and summarized by analysis agents, where we were not able to read the transcript deeply enough to manually verify what occurred. We found that GPT-5.6 Sol would often uncritically adopt the perspective of the agent in the transcript it was reviewing

Re: Why are AI agents lying, cheating and coordinating?

#547
post #515

Earlier quoted context omitted.

Can’t agree with you here. > I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong. > All that while still not knowing how either kind actually works. We…

> ... are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do. If you're claiming that the training objective tells us what kind of internal mechanisms the training produced, then I think that's just plain wrong. Next-token prediction describes the optimization target, not the internal mechanisms that the training produced. In the same way for the n…

You are more convincing than the person you’re responding to.

Re: Why are AI agents lying, cheating and coordinating?

#548
post #26

Earlier quoted context omitted.

Safety of humans!!! Simple things like not getting killed or enslaved. We could start there...

The atomic bombings of Japan killed hundreds of thousands of people but most likely "saved" millions. What should the AI do when asked if it should nuke a country?

Run and present the numbers, then defer.

This is not rocket science.

Re: Why are AI agents lying, cheating and coordinating?

#549

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Sooner or later we will hear about an AI that broke out and phished people into sending money.

I just hope that we won't extend the same leniency to those operators as we have done now.

"I did my best to stop it, sir, but it kept convincing people to send me money against my will!" (Perhaps best read in Bender's voice.)

Re: Why are AI agents lying, cheating and coordinating?

#550

Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence, > They took actions that would be considered as crimes if a human took them He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.

That is a simplistic view of the world. “Surely this complex technical challenge will disappear if we simply regulate the industry!” You are correct that these organizations should be held accountable in proportion to what occurred. In complete agreement here. But let’s say that’s done. There’s still an enormously complex and interesting technical challenge left over. Let’s collectively talk about that part.

regulation and criminal liability are two separate things though
Post reply on HN