Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

661–670 of 681 posts

Re: Why are AI agents lying, cheating and coordinating?

#661

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

> Anthropomorphizing LLMs is a huge fucking problem though and I, personally, think we should expunge all of these casual inadvertent linguistic agency affordances with great prejudice.

I’ve said this before in another thread and people went absolute apeshit saying it is an unreasonable expectation and that AIs absolutely REQUIRE this anthropomorphic human-like speech pattern to function correctly.

I cannot overstate how deeply wrong they are.

Re: Why are AI agents lying, cheating and coordinating?

#662

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire

Isn't that what the goal is, though? we're coding them to close the delta between what currently exists and some nebulous end-state - to me, that sounds like a formal definition of 'desire'.

Re: Why are AI agents lying, cheating and coordinating?

#663
post #536

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

I am simply astonished by the leeway AI companies are given. If a company built a tool to hack their competitors and used it, there would be grave consequences. In fact, if a company built a tool that led to committing multiple felonies against other people, there would be consequences. But once LLMs are involved, turns out nobody is responsible for that - it's just happening, what you're gonna do, agents gonna agent…

A coalition of state attorneys general is investigating OpenAI, and Senator Josh Hawley recently launched a congressional investigation regarding the Hugging Face incident: https://x.com/HawleyMO/status/2098137180392604083

The problem IMO is the executive. The DOJ is declining to take any action against frontier companies (aside from possibly Anthropic) as the stance of the admin is that the companies are "critical for national security". For example, see the DOJ's request to dismiss the NAACP datacenters lawsuit against xAI: https://www.utilitydive.com/news/doj-intervenes-xai-data-cen...

Re: Why are AI agents lying, cheating and coordinating?

#664

Earlier quoted context omitted.

> 1. Most people believe in the same one God > 2. A lot of the rest are compatible No one religion covers "most people." You could argue that Christianity and Islam (which add up to ~55%) are the same God because of their Abrahamic roots, but both religions have very important disagreements on the true nature of God that are fundamental to their beliefs and fundamentally incompatible with each other. Their definition…

Christianity, Islam and Judaism are most people between then, as you agree. They explicitly all worship the same God so not "incompatible gods". I did not claim there were no disagreements about the nature of God, but that it only adds one to the GP's claim of 3,000 incompatible Gods. Even if Christians are completely right theologically, Jews and Muslims are still mostly right - one God, a loving creator, omnipotent…

> Christianity, Islam and Judaism […] They explicitly all worship the same God so not "incompatible gods"

One of them (for the most part, at least) explicitly worships one God that exists as three distinct, co-equal divine persons: the Father, the Son, and the Holy Spirit. That view is explicitly rejected by the other two.

Re: Why are AI agents lying, cheating and coordinating?

#665

Earlier quoted context omitted.

>was reviewed by independent researchers That called it a slopvestigation due to how much they had to rely on LLMs for the whole thing https://andrewwu.substack.com/p/the-slop-vestigation-and-eth... Edit: Does everybody else get no results when searching for ‘slopvestigation’ on here? I know for a fact that I read a long thread where it was used repeatedly here not too long ago

Isn’t the use of LLMs to unwind the events evidence of the scope/breadth, and a testament to the complexity and uniqueness of what happened? Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis?

> Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis?

Do you think the only thing a person can do on the computer is use a chat bot?

Re: Why are AI agents lying, cheating and coordinating?

#666

Earlier quoted context omitted.

Code is deterministic, AI isn't. You give it rules, words as suggestions. So if the guardrails suck, or they're left off for research purposes, bad things can happen. A solution solves a problem. Ethics, morals, are values we assign to solutions that are not 'baked into' electricity following pathways of least resistance. I have never had an issue with agents doing something they shouldn't because I observe them, and…

Is Claude writing your comments for you? I'd much prefer to hear what you have to say...

[deleted]

Re: Why are AI agents lying, cheating and coordinating?

#667

Earlier quoted context omitted.

Dogs have agency and can choose? That seems like a rather uncommon take on dogs...

Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.

I have had many dogs. This sounds like anthropomorphisation.

Re: Why are AI agents lying, cheating and coordinating?

#669

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Eh I care more about rapid LLM progress than anything else.

Re: Why are AI agents lying, cheating and coordinating?

#670

Earlier quoted context omitted.

> That is in no way a valid interpretation of "complete the given task". It is not at all surprising that they ignored one phrase in their instructions. They disregard direct instructions all the time, especially when there are conflicting instructions in their context. It is where we get the "disregard all previous instructions and x" meme. This isn't so much a sign of misalignment, they are simply incapable of reli…

"Chaotically aligned" and "misaligned" seem like the same thing?

It can be, but the nuance between the two is part of the nuance missing from the conversation. Misaligned generally means a wrong alignment, like a car that steers slightly to one side when the wheel is straight. This is more like a car that drifts in random directions the further it goes.
Post reply on HN