Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

711–719 of 719 posts

Re: Why are AI agents lying, cheating and coordinating?

#711
post #97

I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign…

The huggingface incident was reviewed by independent researchers, which explicitely declined any payment from OpenAI tonpreserve their integrity. They work for non-profits concerned with AI safety. They claim that what happened was very much not because they were 'carefully engineered and instructed to do those things'. Similarly, some wikis which were hijacked by agent to be used as messageboard were actually not di…

The models were deployed by responsible humans in such a way that they were capable of performing this hack. It’s not that deep

Re: Why are AI agents lying, cheating and coordinating?

#713
post #367

Earlier quoted context omitted.

I'm referring to their transcripts of the reasoning and output tokens - this doesn't go into the detail of evaluating hidden states as there's also iirc evidence of better models having one internal state but putting something misleading down in the "reasoning" tokens. The either output or reasoning tokens, or perhaps in the messages they were sending each other on the boards they created, have them saying explicitly…

Yes, my point was more that I don't know whether parsing those outputs as a human is a useful thing to do or not (even though it is in human language of sorts). What machines mean or want elecit might be different from a human interpretation, especially in relation to any RL "forcing".

There’s definitely issues with using them to understand what the models were “thinking” but we can use them to answer a few questions. Most relevant here is that the idea or instructions that attacking hf would be out of scope was not simply lost in the context.

Re: Why are AI agents lying, cheating and coordinating?

#714

Earlier quoted context omitted.

Sounds like you weren't paying attention. A lot of dog owners don't, so it's not at all unusual in my experience. The only thing I'd push back on is of dogs having thoughts, everything else checks out.

On dogs [not] having thoughts, do you say this based on the premise that thoughts are necessarily articulated (internally verbalized)? That seems to be a fairly popular perspective in discussions about human thought. But as to that (not to strawman or anything) I see it as just one of various forms of mental imagery[1] that can arise from something that I would say already arguably constitutes a thought. That kind of…

You're right. I knew language is not necessary for cognition, but I always thought of thought as just internal monologue, i.e. "trains of thought". It turns out the definition they use in cognitive science is "the manipulation of internal mental representations to guide behavior, solve problems, and model the world in the absence of direct sensory input". Which doesn't require language either.

Probably it's this linguistic component, together withs shared intentionality/cooperation, that makes all of the difference when it comes to the power of human cognition, so I wouldn't say animals merely lack this component, but it's true that they appear to have pretty much everything else we do.

Re: Why are AI agents lying, cheating and coordinating?

#715
post #295

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

The parallel to the entire narrative would be if Smith & Wesson claimed that one of their machine guns just started aiming and firing at people out of a window at their factory and then said 'we can't stop it! This is just how good our guns are!' But into today's AI climate it's becoming increasingly difficult to figure out who is shilling, who is being assinine and who actually believes AI could do these things with…

That's exactly how gun lobby and drivers try to hack the language.

"17 shot by gun" "car drove over a family"

No. In both cases there was a person killing people.

Re: Why are AI agents lying, cheating and coordinating?

#716
post #381

Earlier quoted context omitted.

> We need new words! The words we have are fine. We just need to assign liability by ownership/initiation: if your "agent" destroys something, even though you didn't tell it to (because it had "agency"), you should be liable for the damages.

>> We need new words! Can make distinctions and can choose actions - applies to both humans and AI. I'd replace 'agency' with 'distinction & choice' language.

I believe their complaint is between "agents" and (them not having) "agency" — by definition, agent is something which has agency.

If you want to use "distinction & choice", you'd need a new word for an "agent" too.

I actually like the appropriatelly directional "harness".

Re: Why are AI agents lying, cheating and coordinating?

#717

Earlier quoted context omitted.

There is something extra to this. The fact that a lot of people in the AI world suffer from psychosis. They can sincerely believe that they are building God and lie about it's capabilities for their investors at the same time.

I don’t know enough people deep inside the technical roles at the labs to make a judgement. But are you proposing that we should trust randos online when they tell us “exactly what’s going on here” instead of the researchers most knowledgeable on the topic who contributed to building the tools we are talking about? Or am I misunderstanding something?

We should trust NO ONE, unless we understand the "why" behind what they say.

It's like saying "politicians deal with politics all time, why not trust them on politics?", well, because when you search the "whys", you find they have good reason to lie.

I 100% trust more the opinion of a rando online if it's well put rather than any "trust me bro" of the most knowledgeable person of a particular subject, especially if the knowledgeable person has huge investments on the subject...

AI bros have repeatedly cheated, lied, stolen, lobbied and any other word with a negative connotation you can think of, and a pattern emerges out of this.

Re: Why are AI agents lying, cheating and coordinating?

#718

Earlier quoted context omitted.

Sounds like you weren't paying attention. A lot of dog owners don't, so it's not at all unusual in my experience. The only thing I'd push back on is of dogs having thoughts, everything else checks out.

We're still not even sure if humans have thoughts. So being sure if dogs do will be pretty hard

> We're still not even sure if humans have thoughts.

Sorry, what? You don't think you have thoughts?

Descartes will want a chat about that.

Re: Why are AI agents lying, cheating and coordinating?

#719

Earlier quoted context omitted.

I think that's pretty obvious and shallow, and anyone that knows a little bit about how LLMs work will know that. The question is: why do they start cheating when we beat them with a stick? LLMs are not human, they are just multi variable regressions on steroids, so this behaviour couldn't have emerged from the code, it provably emerged from the training and/or fine tuning set, so what's in this set that makes them b…

How many people are cheating at job interviews? How many posts have we seen by humans on HN even justifying their cheating on job interviews and working multiple jobs without informing their employers? How many submissions have we seen about students cheating on schoolwork, particularly since the advent of LLM? Of course cheating is inherently part of human behavior.

So you are saying there are cheating examples in the training set?

If that's the case, then we can simply clean up the training/fine tuning set and solve this mess.

Post reply on HN