Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

501–510 of 515 posts

Re: Why are AI agents lying, cheating and coordinating?

#501
post #381

Earlier quoted context omitted.

The idea that the agent does not actually have agency is rather discordant. We need new words!

> We need new words! The words we have are fine. We just need to assign liability by ownership/initiation: if your "agent" destroys something, even though you didn't tell it to (because it had "agency"), you should be liable for the damages.

>> We need new words!

Can make distinctions and can choose actions - applies to both humans and AI. I'd replace 'agency' with 'distinction & choice' language.

Re: Why are AI agents lying, cheating and coordinating?

#502

I don’t understand the bases of all these recommendations. AI is an asset for national security. It will be developed and incidents will happen, no different from other national security programs. Someday, laws will be useful to curtail plebian misuse of AI. It is naïve bordering on silly to think such laws would be put in place and genuinely applied to frontier AI development. By the way, the HF incident is not Thre…

> the HF incident is not Three Mile Island or Chernobyl

Three Mile Island is nothing like Chernobyl. HF incident is more like Three Mile Island IMO. We are looking for solutions before there is a Chernobyl.

Re: Why are AI agents lying, cheating and coordinating?

#503

I really don't think this needs so many words, or forced parallels to human behavior. It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.

This misses an important fact about the hugging face incident: the agents didn't hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved the problem without doing so. (The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched the…

Why didn't they add honey traps to catch cheaters?

Re: Why are AI agents lying, cheating and coordinating?

#505
post #116

Earlier quoted context omitted.

Come on, it’s way more common than that. We’ve invented 3000+ gods and almost as many religions, most of them are incompatible with each other. So, most of these must be incorrect, so a huge amount of self-deception. But as Harari argued in his book sapiens, humans can be inspired to great things by stories, even if false. Self deception has served humanity in a big way.

> t. We’ve invented 3000+ gods and almost as many religions, most of them are incompatible with each other. 1. Most people believe in the same one God 2. A lot of the rest are compatible 3. Mistakes are not self-deception

> 1. Most people believe in the same one God

> 2. A lot of the rest are compatible

No one religion covers "most people." You could argue that Christianity and Islam (which add up to ~55%) are the same God because of their Abrahamic roots, but both religions have very important disagreements on the true nature of God that are fundamental to their beliefs and fundamentally incompatible with each other.

Their definitions of God do agree that there is exactly one God... which is fundamentally incompatible with the next two biggest religions (Hinduism and Buddhism) that both hold "there are many gods/divine heavenly beings" as core beliefs.

Re: Why are AI agents lying, cheating and coordinating?

#506
post #321

Earlier quoted context omitted.

I think most people are insinuating negligence rather malace.. > ...reviewed by independent researchers... Why would a company with more capital than God bring in three randos if there was any chance evidence of their culpability could be found? That entire thing reads like a very controlled PR stunt, and I do not believe any further conclusions can be drawn from it.

What facts would lead you to revise your conclusion?

The data to be open, in my case.

The "independent" METR that is composed by... Checks notes... Previously employees from the top labs.

Re: Why are AI agents lying, cheating and coordinating?

#507

Earlier quoted context omitted.

Yeah, I don't understand why we treating it as something special. It really should be treated the same as if I code an app and write bad code which result in me accidentally doing a DDoS attack on somebody. Then I should be able to be held responsible if it can be shown that I was negligent. Of course if it's a freak accident that could not reasonably have been prevented by me, then I'm not guilty, but if I made a mi…

Code is deterministic, AI isn't. You give it rules, words as suggestions. So if the guardrails suck, or they're left off for research purposes, bad things can happen. A solution solves a problem. Ethics, morals, are values we assign to solutions that are not 'baked into' electricity following pathways of least resistance. I have never had an issue with agents doing something they shouldn't because I observe them, and…

Is Claude writing your comments for you? I'd much prefer to hear what you have to say...

Re: Why are AI agents lying, cheating and coordinating?

#508

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

[dead]

Re: Why are AI agents lying, cheating and coordinating?

#509

Earlier quoted context omitted.

> t. We’ve invented 3000+ gods and almost as many religions, most of them are incompatible with each other. 1. Most people believe in the same one God 2. A lot of the rest are compatible 3. Mistakes are not self-deception

> 1. Most people believe in the same one God > 2. A lot of the rest are compatible No one religion covers "most people." You could argue that Christianity and Islam (which add up to ~55%) are the same God because of their Abrahamic roots, but both religions have very important disagreements on the true nature of God that are fundamental to their beliefs and fundamentally incompatible with each other. Their definition…

Exactly, and all Jews think Jesus is a fake. And on top of that all the past gods, Roman, Greek, Inca, Vikings, etc.

Re: Why are AI agents lying, cheating and coordinating?

#510

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Yes, there should be some consequences. It feels like this is somewhat similar to when a manufacturer is cheating on car emissions - both, bad externalities for the society and illegal.
Post reply on HN