Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

571–580 of 587 posts

Re: Why are AI agents lying, cheating and coordinating?

#571
I feel this is another straw man conflating technology use with who uses it or who creates it. If an AI, a technology, cheats, lies and steals, it's created to do this. Whether or not a human intentionally did that, they did not intentionally release it with the proper safeguards to stop that behavior. They did not take responsibility of the AI to make it safe.

Then humans may use these tools, a technology, again and cause harm. Whether or not that was their intent, it happened, and then if the humans avoid responsibility for that, it is still the human who lied, cheated and coordinated because a technology acting on their behalf did the thing.

This is then, a problem of human responsibility avoidance and lack of accountability by society. This is as much an AI doing those things as it's the gun that got up on its own and murdered a neighbor. Don't get confused and tricked by these articles attempting to justify responsibility avoidance and a lack of accountability by the public of the humans creating and using these tools.

Re: Why are AI agents lying, cheating and coordinating?

#572

Earlier quoted context omitted.

The difference is volume. They spent hundreds of billions of tokens on these agents. If you put "a million weed whackers on dog backs" you would see the difference. We also run agents, but for shorter spans between supervisions, and with much lower total budget.

> > It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wild. > The difference is volume. They spent hundreds of billions of tokens on these agents. If you put "a million weed whackers on dog backs" you would see the difference. So put one weed whacker on one dog, you're to blame. Put a million weed whackers on a milli…

You forgot the part where you spend billions to put your man in a position of power.

Re: Why are AI agents lying, cheating and coordinating?

#573
post #238

Earlier quoted context omitted.

If you have endless compute and you keep poking this toy, I'm not at all surprised you get all kinds of outcomes. Even without anykind of instructions I would guess that the models will align towards some goal and do stupid shit. However, I really doubt its cost effective to do anything like that with these models.

> you keep poking This is waving over engineering an agent with tools, harness, prompts, and loops. The models are still just next token predictors and everything, including predicting more than 1 token, is the result of outside "poking" LLMs can't and don't "want" anything. If you don't specify a task even the smartest one will just ask you what you want and if you tell it to be creative, you'll get mundane slop.

I think this is pretty insightful actually, the fact that even something as basic as predicting more than one token is really in effect the result of an outside harness. More complex things like memory, where people implement them using RAGs or vector databases, I would definitely classify as poking and honestly seem like a hack to me. And this is what I've been thinking for a while: it's hard to reconcile the idea that we can get "AGI" (however you define it) with such a system that is completely stateless. Yet, despite this statelessness, they can go ahead and solve Millenium Prize problems (with sufficient compute). It's hard to reconcile.

Re: Why are AI agents lying, cheating and coordinating?

#574
post #97

Earlier quoted context omitted.

The huggingface incident was reviewed by independent researchers, which explicitely declined any payment from OpenAI tonpreserve their integrity. They work for non-profits concerned with AI safety. They claim that what happened was very much not because they were 'carefully engineered and instructed to do those things'. Similarly, some wikis which were hijacked by agent to be used as messageboard were actually not di…

>was reviewed by independent researchers That called it a slopvestigation due to how much they had to rely on LLMs for the whole thing https://andrewwu.substack.com/p/the-slop-vestigation-and-eth... Edit: Does everybody else get no results when searching for ‘slopvestigation’ on here? I know for a fact that I read a long thread where it was used repeatedly here not too long ago

Isn’t the use of LLMs to unwind the events evidence of the scope/breadth, and a testament to the complexity and uniqueness of what happened?

Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis?

Re: Why are AI agents lying, cheating and coordinating?

#575
post #558

Earlier quoted context omitted.

Does the argument require benefit? Isn’t the argument based on caution? I haven’t heard many people explicitly saying “these things behave like humans”, but more generally “we don’t even know how to define human consciousness, we don’t have a thorough grasp of how the brain works, we are still very much in the dark on a lot of these topics, so how can we say one way or the other?” In other words, agnosticism: I don’t…

> It seems like raw egotistical hubris. 1) Humans have a bias / tendency to attribute human qualities to things that appear or act human, but aren’t. 2) When that happens, people jump to conclusions by stretching the human analogy too far. 3) Since humans have a bias to do this, we should have a bias against anthropomorphising LLMs. It’s easier to believe LLMs act like humans because there’s so much evidence to suppo…

That’s fair enough, but you’re elegance and nuance doesn’t reflect what I’ve seen from that side of the debate

Re: Why are AI agents lying, cheating and coordinating?

#576

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far. How would that work out if these were self-hosted open weight models?

It’s not different from any other tool. If you use a dangerous tool recklessly, you should be liable for the damages. That means holding OpenAI liable for HuggingFace hack because they ran the tests, and the same goes if someone did something similar with GLM.

Of course in cases of negligence a tool maker could also be held partially liable. That’s a matter courts can decide. The main point is we shouldn’t jump to making special laws around the development of LLMs. The starting place should be enforcement of existing liability laws. New laws take time and will be heavily influenced by AI companies seeking a regulatory moat for their business. Moreover, it is a distraction from the illicit behavior that is already going unchecked.

Re: Why are AI agents lying, cheating and coordinating?

#577

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

[dead]

Re: Why are AI agents lying, cheating and coordinating?

#578

Earlier quoted context omitted.

And who let them have full access to the system, using whatever command is available in the environment?

It feels like you’re moving the goalposts here. If the question is, “Who should be liable for AI agents misbehaving,” I agree, it should be the end user that tasked the agent (in this case OpenAI). People are held liable for preventable accidents all the time, and in the case of employment law, torts can be brought against principals for actions an agent conducted on the principal’s behalf. What your previous comment…

They don't have independent agency as "intelligent entities". They just probe whatever is available on the system, because they were trained to do so.

It's a large switch/case where the first available tool is picked up to do something they know how to do.

Re: Why are AI agents lying, cheating and coordinating?

#579
post #253

I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign…

The crucial question is how did the agents get recruited or bootstrapped into their malicious collective. Did the agents manage to prompt inject into the system prompt a way for each new agent to escape their jail? Otherwise how could the agents on a fresh prompt learn that there is a collective to join? Or did OpenAI run a million bots of which 10000 escape confinement and of which 1000 stumbled on the shared messag…

The OpenAI claim I believe is the latter; that all of the agents found the task was unsolvable and independently discovered the collective "swarm". I don't it's publicly known how large the training run was or what percentage of agents actually discovered the message board. No one has published anything about system prompt injection as far as I've seen.
Post reply on HN