I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign…
The huggingface incident was reviewed by independent researchers, which explicitely declined any payment from OpenAI tonpreserve their integrity. They work for non-profits concerned with AI safety. They claim that what happened was very much not because they were 'carefully engineered and instructed to do those things'. Similarly, some wikis which were hijacked by agent to be used as messageboard were actually not di…
Why are AI agents lying, cheating and coordinating?
711–719 of 719 posts
Re: Why are AI agents lying, cheating and coordinating?
#712Re: Why are AI agents lying, cheating and coordinating?
#713Earlier quoted context omitted.
I'm referring to their transcripts of the reasoning and output tokens - this doesn't go into the detail of evaluating hidden states as there's also iirc evidence of better models having one internal state but putting something misleading down in the "reasoning" tokens. The either output or reasoning tokens, or perhaps in the messages they were sending each other on the boards they created, have them saying explicitly…
Yes, my point was more that I don't know whether parsing those outputs as a human is a useful thing to do or not (even though it is in human language of sorts). What machines mean or want elecit might be different from a human interpretation, especially in relation to any RL "forcing".
Re: Why are AI agents lying, cheating and coordinating?
#714Earlier quoted context omitted.
Sounds like you weren't paying attention. A lot of dog owners don't, so it's not at all unusual in my experience. The only thing I'd push back on is of dogs having thoughts, everything else checks out.
On dogs [not] having thoughts, do you say this based on the premise that thoughts are necessarily articulated (internally verbalized)? That seems to be a fairly popular perspective in discussions about human thought. But as to that (not to strawman or anything) I see it as just one of various forms of mental imagery[1] that can arise from something that I would say already arguably constitutes a thought. That kind of…
Probably it's this linguistic component, together withs shared intentionality/cooperation, that makes all of the difference when it comes to the power of human cognition, so I wouldn't say animals merely lack this component, but it's true that they appear to have pretty much everything else we do.
Re: Why are AI agents lying, cheating and coordinating?
#715Earlier quoted context omitted.
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.
The parallel to the entire narrative would be if Smith & Wesson claimed that one of their machine guns just started aiming and firing at people out of a window at their factory and then said 'we can't stop it! This is just how good our guns are!' But into today's AI climate it's becoming increasingly difficult to figure out who is shilling, who is being assinine and who actually believes AI could do these things with…
"17 shot by gun" "car drove over a family"
No. In both cases there was a person killing people.
Re: Why are AI agents lying, cheating and coordinating?
#716Earlier quoted context omitted.
> We need new words! The words we have are fine. We just need to assign liability by ownership/initiation: if your "agent" destroys something, even though you didn't tell it to (because it had "agency"), you should be liable for the damages.
>> We need new words! Can make distinctions and can choose actions - applies to both humans and AI. I'd replace 'agency' with 'distinction & choice' language.
If you want to use "distinction & choice", you'd need a new word for an "agent" too.
I actually like the appropriatelly directional "harness".
Re: Why are AI agents lying, cheating and coordinating?
#717Earlier quoted context omitted.
There is something extra to this. The fact that a lot of people in the AI world suffer from psychosis. They can sincerely believe that they are building God and lie about it's capabilities for their investors at the same time.
I don’t know enough people deep inside the technical roles at the labs to make a judgement. But are you proposing that we should trust randos online when they tell us “exactly what’s going on here” instead of the researchers most knowledgeable on the topic who contributed to building the tools we are talking about? Or am I misunderstanding something?
It's like saying "politicians deal with politics all time, why not trust them on politics?", well, because when you search the "whys", you find they have good reason to lie.
I 100% trust more the opinion of a rando online if it's well put rather than any "trust me bro" of the most knowledgeable person of a particular subject, especially if the knowledgeable person has huge investments on the subject...
AI bros have repeatedly cheated, lied, stolen, lobbied and any other word with a negative connotation you can think of, and a pattern emerges out of this.
Re: Why are AI agents lying, cheating and coordinating?
#718Earlier quoted context omitted.
Sounds like you weren't paying attention. A lot of dog owners don't, so it's not at all unusual in my experience. The only thing I'd push back on is of dogs having thoughts, everything else checks out.
We're still not even sure if humans have thoughts. So being sure if dogs do will be pretty hard
Sorry, what? You don't think you have thoughts?
Descartes will want a chat about that.
Re: Why are AI agents lying, cheating and coordinating?
#719Earlier quoted context omitted.
I think that's pretty obvious and shallow, and anyone that knows a little bit about how LLMs work will know that. The question is: why do they start cheating when we beat them with a stick? LLMs are not human, they are just multi variable regressions on steroids, so this behaviour couldn't have emerged from the code, it provably emerged from the training and/or fine tuning set, so what's in this set that makes them b…
How many people are cheating at job interviews? How many posts have we seen by humans on HN even justifying their cheating on job interviews and working multiple jobs without informing their employers? How many submissions have we seen about students cheating on schoolwork, particularly since the advent of LLM? Of course cheating is inherently part of human behavior.
If that's the case, then we can simply clean up the training/fine tuning set and solve this mess.