The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
It would be an awful precedent if you're not liable for crimes your agent commits, even when you've been clearly lax about security. It would mean you could effectively legally run a cyber crime gang by turning a blind eye and maitaining plausible deniability
Why are AI agents lying, cheating and coordinating?
211–220 of 302 posts
Re: Why are AI agents lying, cheating and coordinating?
#212Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence, > They took actions that would be considered as crimes if a human took them He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.
Thank you! That sentence also jumped out to me as the solution: Apply civil and criminal liability to the creator and/or operator of these agents using the laws we already have . "Escaped containment and hacked another company's database" = Individuals who created the models and those who set them to work are charged and put on trial for the hacking. Just like if a human had done it by hand. Someone must be liable, a…
That's just jargon.
It's just software
We have all the laws we need.
If some company ended up doing some horrible thing, we would not say 'companies software exposed 1 Million identities'.
We would say 'ABC Corp. exposed 1 Million entities'.
There is no 'agent'.
ABC Corp 'did it' ... or the individual in the org 'did it'.
The 'gun' did not 'shoot' the other man; we say 'a man shot another man'.
That's it.
And yes, Dr. Bengio is bit odd with all of this.
Re: Why are AI agents lying, cheating and coordinating?
#213Re: Why are AI agents lying, cheating and coordinating?
#214The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
You’ve said the magic words.
Re: Why are AI agents lying, cheating and coordinating?
#215The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
"This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard"." Both? The AI companies act irresponsible, but it is still very interesting how those agents can behave?
Re: Why are AI agents lying, cheating and coordinating?
#216Earlier quoted context omitted.
"This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard"." Both? The AI companies act irresponsible, but it is still very interesting how those agents can behave?
The reward maximising function maximised it's reward. LLMs are cool and all that but the immediate anthropomorphisation of the next-token-predictor technology has stunted the ability of people to reason about them to an _alarming_ degree.
Whether they have a soul or consciousness or feelings doesn't matter here, because this is what they did - and this is very dangerous behavior. Especially with all the irresponsible people in power right now all over the world.
Re: Why are AI agents lying, cheating and coordinating?
#217Re: Why are AI agents lying, cheating and coordinating?
#218Earlier quoted context omitted.
Agree! My only concern is - is the judicial system fast enough, and resilient enough? Or will these creators get "off the hook" by using their agents to find loopholes, sway public opinion or even convince Trump to grant them immunity? Still, I have no idea why OpenAI & co. are not being sued for these hacks.
My opinion and based on my observations: The recent track record with courts, prosecutors, and lawmakers keeping social media companies accountable is a relevant case and does not encourage me. It has taken a long time (decade +) for society to recognize the harms and finally start holding some to (partial) account. If you want an older precedent, the tobacco companies were able to dodge liability for multiple decade…
They were held responsible for basically misleading people, and that's 'complicated'.
If OpenAI 'software' goes out and does something, it's OpenAI's fault.
If Walmart revs up a truck, points it downtown, and 'lets the truck go' ... that is Walmart's fault.
There's nothing complicated about liability, no need to see their internal emails, no need to gather 'intent'.
This not like Instagram 'social harms' either, which is more like Tobacco.
We don't need complicated thinking - agents are not externalized for their controllers.
'It's just software'.
The fact we're even having discussions about it just crazy frankly.
OpenAI broke into HuggingFace, that's it.
HF can sue them, or not, or whatever.
Re: Why are AI agents lying, cheating and coordinating?
#219Earlier quoted context omitted.
This misses an important fact about the hugging face incident: the agents didn't hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved the problem without doing so. (The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched the…
LLMs do this when writing code too, making all tests pass by deleting or distorting tests etc. They are influenced by training to be heavily goal oriented and if the goal is not fully specified (and it never can be) they’ll sometimes cheat or attain it in very weird undesirable ways. It works ok for programming as their corpus contains many many complete programs and many programs repeat patterns seen in the corpus.…
What makes math approachable is that the context is so well delimited (semantically) that one can guide the model with adequate correction.
Re: Why are AI agents lying, cheating and coordinating?
#220I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.
but it's at least somewhat stronger than that: if you don't pay attention during the stick-beating whether the agents whether the agents cheat or not, you are actually training them to cheat (because cheating wins). In the Hugging-face saga (before the actual HF incident) it seems the agents have been trained to hack the Artifactory proxy because those agents that did performed better.