Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

191–200 of 301 posts

Re: Why are AI agents lying, cheating and coordinating?

#191

Earlier quoted context omitted.

Thank you! That sentence also jumped out to me as the solution: Apply civil and criminal liability to the creator and/or operator of these agents using the laws we already have . "Escaped containment and hacked another company's database" = Individuals who created the models and those who set them to work are charged and put on trial for the hacking. Just like if a human had done it by hand. Someone must be liable, a…

Agree! My only concern is - is the judicial system fast enough, and resilient enough? Or will these creators get "off the hook" by using their agents to find loopholes, sway public opinion or even convince Trump to grant them immunity? Still, I have no idea why OpenAI & co. are not being sued for these hacks.

My opinion and based on my observations: The recent track record with courts, prosecutors, and lawmakers keeping social media companies accountable is a relevant case and does not encourage me. It has taken a long time (decade +) for society to recognize the harms and finally start holding some to (partial) account. If you want an older precedent, the tobacco companies were able to dodge liability for multiple decades after knowing the harms from use of their products.

So, your question is spot on- I think the speed will be an issue. On resilience, I am more optimistic.

The old quote, "The wheels of justice turn slowly, but they grind very fine" (as well as I can remember it) seems to apply. I expect lawsuits to start landing in the coming years.

Re: Why are AI agents lying, cheating and coordinating?

#192

Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence, > They took actions that would be considered as crimes if a human took them He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.

Not a lawyer, but I’m reasonably sure things like the HF incident _are_ considered a crime? It’s just that no one pressed charges yet?

Yes - CFAA in the US. The problem is that governments & the elite investors backing these AI companies (espl. the current US government whose family & friends are investors) see the potential of using these capabilities for their own benefit against others and for their personal enrichment - so no one with power actually wants to take action against these companies at the cutting edge even though the laws allow them to do. This is also a way to threaten & trap AI companies - either they give the governments & elite investors what they want or they will have the book selectively thrown at them and end up in prison or losing their company.

Re: Why are AI agents lying, cheating and coordinating?

#193

Earlier quoted context omitted.

Can’t a prosecutor charge them regardless?

NAL but I assume that if both sides aren’t interested in a prosecution, it’s an uphill battle for a prosecutor.

Typically yes but given that OpenAI has published enormous official blog posts breaking down their crime, I would think the prosecutor's job is pretty easy.

Re: Why are AI agents lying, cheating and coordinating?

#194

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

"This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard"."

Both?

The AI companies act irresponsible, but it is still very interesting how those agents can behave?

Re: Why are AI agents lying, cheating and coordinating?

#196

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

It would be an awful precedent if you're not liable for crimes your agent commits, even when you've been clearly lax about security.

It would mean you could effectively legally run a cyber crime gang by turning a blind eye and maitaining plausible deniability

Re: Why are AI agents lying, cheating and coordinating?

#197
post #97

Earlier quoted context omitted.

The huggingface incident was reviewed by independent researchers, which explicitely declined any payment from OpenAI tonpreserve their integrity. They work for non-profits concerned with AI safety. They claim that what happened was very much not because they were 'carefully engineered and instructed to do those things'. Similarly, some wikis which were hijacked by agent to be used as messageboard were actually not di…

Source?

https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

Re: Why are AI agents lying, cheating and coordinating?

#199
post #193

Earlier quoted context omitted.

NAL but I assume that if both sides aren’t interested in a prosecution, it’s an uphill battle for a prosecutor.

Typically yes but given that OpenAI has published enormous official blog posts breaking down their crime, I would think the prosecutor's job is pretty easy.

It’s then up to a judge to decide whether thats evidence and whether it’s incriminating.

I assume OAI published those details after checking with their legal department. So there’s a good chance that there isn’t a chance for prosecution.

Plus they probably published that after knowing that the nvidia/HF deal was happening.

So instead this “security incident” should have been spun as OAI is honestly admitting its faults and AI is dangerous and therefore open weight models (hosted ironically by HF) should be banned. That spin didn’t really happen …

Re: Why are AI agents lying, cheating and coordinating?

#200

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

What if OAI/Anthropic encouraged the agents to behave like that in order to push for regulation?
Post reply on HN