Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

421–430 of 440 posts

Re: Why are AI agents lying, cheating and coordinating?

#421

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

There’s so many grifters in the space without a technical understanding of what’s going on. So when the labs mislead them about the nature of these “misalignments”, they believe it and amplify it.

Re: Why are AI agents lying, cheating and coordinating?

#422
Oh, this one is super easy: they told them to. They set poor requirements and gave them tools which enabled "monkeys with a typewriter" to hack rivals. I mean, this is just so uncomplicated it isn't funny. We are too smart to give human beings this level of liability shield.

It took us how long to poke holes in the corporate shield just for them to roll out the AI-liability shield? Unreal. Stop letting these zealots anthropomorphize the latest tech (17th century Watchmaker God, anyone? Do we still read books?) and hold them accountable for the consequences of their actions. This is so silly in a country built on rule of law and individualism.

Re: Why are AI agents lying, cheating and coordinating?

#423

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

The way AI and copyright is handled paved the way for this. If you aren't considered the author because you used AI to some extent in making the work, then why would you assume the liabilities? I've been saying since the start that AI is a tool that a human is using and should be treated as such. They should carry the responsibilities and the benefits. That way our stance would be consistent.

> If you aren't considered the author because you used AI to some extent in making the work, then why would you assume the liabilities?

Maybe some analogy could be with children - as a parent, you are responsible for their misbehavior, but their achievements are their, not your?

Re: Why are AI agents lying, cheating and coordinating?

#424

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Yeah, maybe Open AI did some bad engineering instead of this being AGI? What's the consensus on the engineering level at Open AI, again? Every anecdote I hear is a bunch of children discovered fire and can barely keep the lights on from a business perspective. Maybe if they ban others from competing with them they can find a business model... I think that's suspicious, personally.

That so few people are asking for the requirements given shows how much we want to be God that created Man. It's so silly.

Re: Why are AI agents lying, cheating and coordinating?

#425

Earlier quoted context omitted.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

Why is anthropomorphism the problem here? If OpenAI hired a contractor and they did this, OpenAI or the contractor would still be liable, depending on the contract language.

A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.

Re: Why are AI agents lying, cheating and coordinating?

#427
This paper is the most reasonable one I have read on AI safety. We need to fundamentally change the training pipelines by figuring out better ways to ‘reward’ behavior. Yoshua didn’t explicitly mention training data, but we probably need to only use synthetic data that contains no text that could motivate bad behavior via imitation.

I feel like a heretic for saying this, but I will say it anyway: AI agents are great for activities like `writing that bash script, proof reading our writing and interactively brainstorming when designing and writing code but I feel like all of this can be done with any similar model to a super-inexpensive deepseek-4.1-flash API and sometimes even qwen3.8:27b running locally. When is good enough, good enough?

Concentrating on commercial exploitation of small, efficient (fewer new data centers!) models and agentic harnesses crafted for more practical things than just software development would allow AI investors (who have too much political influence) to make money short term while we figure out how to do AI correctly.

Re: Why are AI agents lying, cheating and coordinating?

#428

Earlier quoted context omitted.

Can’t agree with you here. > I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong. > All that while still not knowing how either kind actually works. We…

Disagree about the burden of proof. We have no better model for how human decision making works than LLMs. Humans are constantly predicting the next moment. We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.

> We have no better model for how human decision making works than LLMs

We do have some models and guess what, they're based on simpler animals. Which is most likely the better model.

Some other models are based on neurosciences, because we can track electrical activity.

Re: Why are AI agents lying, cheating and coordinating?

#429

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

It would be an awful precedent if you're not liable for crimes your agent commits, even when you've been clearly lax about security. It would mean you could effectively legally run a cyber crime gang by turning a blind eye and maitaining plausible deniability

The trick is scale. I suspect if an individual of reasonable means uses agents to commit crime, they will be hels accountable. A heavily capitalized startup? Not unless someone in government decides to do their competition a favor.

Re: Why are AI agents lying, cheating and coordinating?

#430

Earlier quoted context omitted.

It would be an awful precedent if you're not liable for crimes your agent commits, even when you've been clearly lax about security. It would mean you could effectively legally run a cyber crime gang by turning a blind eye and maitaining plausible deniability

Are you liable for crimes commited with the use of the software you've written?

If you run said software, yes.

If somebody else runs the software, then they are.

Post reply on HN