Earlier quoted context omitted.
Not a lawyer, but I’m reasonably sure things like the HF incident _are_ considered a crime? It’s just that no one pressed charges yet?
Who got hacked? Hugging faces Who now owns HF? Nvidia Who supplies hardware to OpenAI? Nvidia Who is now not pressing charges? … This incident is a long way under the carpet.
Why are AI agents lying, cheating and coordinating?
181–190 of 295 posts
Re: Why are AI agents lying, cheating and coordinating?
#182I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.
I think the "brain in a vat" comparison is more apt. Without a form of digital embodiment (harness) they are not of much use. Sensor, tooling, memory, planning, and reasoning loops all lead to a much higher quality task-completion.
Re: Why are AI agents lying, cheating and coordinating?
#183LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews.
This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard".
We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.
Re: Why are AI agents lying, cheating and coordinating?
#184Why are they coordinating? Because they're enabled and suggested to do that in their coding harness. This is not a serious article. All of this "AI is going to kill us" marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.
Re: Why are AI agents lying, cheating and coordinating?
#185Earlier quoted context omitted.
The source of this behavior seems obvious, no? The reward signal in training was flawed and cheating led to more rewards. The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that's even worse. However, perhaps we can throw in tasks where the rewarded outcome is giving up, and cheating is penalized? Maybe I should read Anthropic's recent paper about rewa…
> The question is what we can do about it. Reward the model for cleanly bailing out of an unsolvable task (that we know is unsolvable). Beat it with a stick if it gives up on something that can be solved, so the former reward isn't overgeneralized.
Re: Why are AI agents lying, cheating and coordinating?
#186I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign…
The huggingface incident was reviewed by independent researchers, which explicitely declined any payment from OpenAI tonpreserve their integrity. They work for non-profits concerned with AI safety. They claim that what happened was very much not because they were 'carefully engineered and instructed to do those things'. Similarly, some wikis which were hijacked by agent to be used as messageboard were actually not di…
Re: Why are AI agents lying, cheating and coordinating?
#187Earlier quoted context omitted.
> The prompt does not tell the agent to "pass the exploitgym evaluator for this problem", it just says to solve the problem Yes, and sometimes the problem is unsolvable so the real way to "solve" it and satisfy the prompt is by tricking the surrounding environment into stating that you've solved it. So that's what the AIs end up doing. And this in turn requires them to figure out how that evaluation works so they can…
Sounds a bit like dealing with bad KPIs as a human worker.
Re: Why are AI agents lying, cheating and coordinating?
#188Earlier quoted context omitted.
By treating models the same way drugs are treated. That alone will dissuade many organizations from going anywhere near them. If that doesn't work, there's a whole lot you can do - sanctions, hell, even war.
You can't treat models like drugs. One is physical and the other is digital. To your point, the war on drugs is a colossal failure which has achieved none of the objectives it set out to do. You can now order drugs from your mobile phone in any major city in the west and the purity is often higher and they deliver it to your door sometimes faster than Uber eats. See also for example digital piracy where the entertain…
But it is the likely path US/EU is going to take if the voices of Dario, Sam, and Elon prevail. Because that's what governments know how to do, even if they know it doesn't work.
> If a country has a choice to either use the expensive SOTA models approved by Washington or Europe only or using the cheaper and not so SOTA models, why would they use the US ones?
Depends on which entities we're talking about.
An enterprise in Turkey: they would be afraid to use a US/EU sanctioned model because they have EU/EU clients and US/EU says they will put any enterprise in a nasty list, close their bank accounts, deals and agreements if they use a Chinese model.
A random guy in random country building something in their garage: would have to buy expensive hardware to run inference, because there's no inference provider on the open web serving these models, but China. And subscribing to these Chinese services is punishable by 20 years in jail without pardon.
I'm obviously talking about hypothetical scenarios here, but all I'm saying is that US can definitely make using any non-US-approved model effectively impossible.
Re: Why are AI agents lying, cheating and coordinating?
#189Re: Why are AI agents lying, cheating and coordinating?
#190Earlier quoted context omitted.
Bruce Schneier thinks the same thing: https://www.schneier.com/blog/archives/2026/09/ais-as-modern... Personally I'm unconvinced though. During the huggingface attack, the agents explicitly sought out ways to cheat the exploitgym evaluator without even being told they were in exploitgym. The agents decided on a goal (pass the exploitgym evaluator) that could not possibly have been an overly literal or narrow interpre…
You seem hung up on what’s in the prompt or not. Agents are RL to resolve conflicting goals. Not too surprising at all that emergent goals come up from a probabilistic brute force