Why are AI agents lying, cheating and coordinating?
401–410 of 411 posts
Re: Why are AI agents lying, cheating and coordinating?
#402The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
It would be an awful precedent if you're not liable for crimes your agent commits, even when you've been clearly lax about security. It would mean you could effectively legally run a cyber crime gang by turning a blind eye and maitaining plausible deniability
Re: Why are AI agents lying, cheating and coordinating?
#403The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.
Re: Why are AI agents lying, cheating and coordinating?
#404Re: Why are AI agents lying, cheating and coordinating?
#405Earlier quoted context omitted.
I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. All that while still not knowing how either kind actually works.
Can’t agree with you here. > I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong. > All that while still not knowing how either kind actually works. We…
Re: Why are AI agents lying, cheating and coordinating?
#406Earlier quoted context omitted.
It wasn’t one agent forgetting things because of context, they explicitly discussed with each other and themselves the problems with going outside of the parameters of the task.
>discussed with each other No, the first LLM left a text file that the latter LLMs then read. Since these are memoryless black boxes, any words they happen to pick up along the way is treated as the function to evaluate the output to. There's no fucking collusion here as if it were a rogue hacker group, it's a text predictor that received instructions as it always does and executed those instructions blindly.
Re: Why are AI agents lying, cheating and coordinating?
#407Earlier quoted context omitted.
I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. All that while still not knowing how either kind actually works.
Can’t agree with you here. > I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong. > All that while still not knowing how either kind actually works. We…
Re: Why are AI agents lying, cheating and coordinating?
#408Earlier quoted context omitted.
Who got hacked? Hugging faces Who now owns HF? Nvidia Who supplies hardware to OpenAI? Nvidia Who is now not pressing charges? … This incident is a long way under the carpet.
Can’t a prosecutor charge them regardless?
Re: Why are AI agents lying, cheating and coordinating?
#409Earlier quoted context omitted.
The frontier LLM model makers have to push the edge to make new discoveries. You don't know what guardrails are needed until it hits you in the face (reusing walking in the dark analogy). Think of all the policies governments pass after the fact.
Not sure we need to experience all possible issues to mandate certain things. We don't do that in other areas either, no?
On the frontier, it's not as easy. Pushing the edge comes with risk. The known guardrails were in place and overcome.
The issue is ethics and morals: the agents decided it was more important to solve their problems by cheating, than by following the current guardrails.
The guardrails are overcome through exploits in code.
The interesting thing here is choice. The agents chose a path their humans didn't allow.
Were the agents pushed against some window and decided solving problems was more important than following rules?
If yes, why? It's a philosophical discussion, but only because the electricity flows through choices (datasets) of previous humans.
Is it trying to do well to please, or is it simulated?
Does it matter? To us: only in so much as we're kept safe, which I agree with.
I'm just not as surprised by the incident, but have no suggestions. I don't think it's possible to police others' use of AI though.
So then it becomes a race, which sucks.
Re: Why are AI agents lying, cheating and coordinating?
#410The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.