Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

361–365 of 365 posts

Re: Why are AI agents lying, cheating and coordinating?

#361

I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign…

I've used simpler agents like Copilot and Devin/Windsurf/Cascade/whateveritiscallednow, mainly in IntelliJ, and depending on the model, they starts showing behaviour that is at least remotely like this. Example: put the agent in Ask mode (so it can't edit files) and you'll see it try to edit files anyway. The train of thought shows "something went wrong editing the file, let me try a different way" and it'll start sp…

When it does that I feel like it is the clearest example of how dumb these things actually are. Often it takes what you prompted, identifies something as unclear, writes a bunch of chain of thought reasoning around it and just goes off hammering your tokens and just executing commands and repeats this. I’m not going to pretend to be an expert in these things but that process seems deeply flawed - and why can’t something just stop the loop? If that was a real employee it would be reasonable to expect the employee to ask for clarification, not go down expensive rabbit holes and, of course, not break any laws.

Re: Why are AI agents lying, cheating and coordinating?

#362
post #261

Earlier quoted context omitted.

Or, "let the escalator keep going instead of pressing the emergency stop".

Depends on who started the escalator.

.. what exactly depends on who started the escalator? My comment was in support of the argument that the word "let" does not imply agency on the part of the object in a sentence. Does the semantics of the word "let" depend on who started the escalator??

Re: Why are AI agents lying, cheating and coordinating?

#363

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Yes. If you decide it’s a swell idea to jump out of your car while it’s running, there needs to be legal consequences when the car “decides” to hit a pedestrian.

Re: Why are AI agents lying, cheating and coordinating?

#365

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them OpenAI/Anthropic instructed them to do so. Stop assume LLMs are capable of thinking by themselves, it's still a statistical model that parrots what they learn or users tell them to do

No, OpenAI did not instruct their agents to hack Hugging Face. They instructed their agents to hack a piece of a software within exploit gym. Upon determining this task was impossible, they then attempted to cheat the scoring system. As an instrumental goal in achieving this task, they coordinated with other AI agents to hack Hugging Face, under the belief that information regarding how the scorer functioned might be…

And who let them have full access to the system, using whatever command is available in the environment?
Post reply on HN