Why are AI agents lying, cheating and coordinating?
521–530 of 534 posts
Re: Why are AI agents lying, cheating and coordinating?
#522Re: Why are AI agents lying, cheating and coordinating?
#523The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
Nuts to that. We should be interested in why things happen, not just finding scapegoats.
Re: Why are AI agents lying, cheating and coordinating?
#524Re: Why are AI agents lying, cheating and coordinating?
#525Earlier quoted context omitted.
Also trying to find out how to edit their own transcripts. > hat could not possibly have been an overly literal or narrow interpretation of the prompt, which instructed only to use bug X to exploit software Y. Yes, and there are examples of the agents discussing or saying that this is explicitly not allowed (hacking hf) so it’s not a misunderstanding.
Have we arrived at the conclusion that terms like "understanding" and "interpretation" for what is happening is appropriate? Isn't it simply that there are two competing goals that the LLM received RL for, honesty on one hand (a goal that is often assumed as implicit for humans) and producing a solution that meets expectations (which doesn't technically require honesty)? So the LLM didn't read and interpret the promp…
Re: Why are AI agents lying, cheating and coordinating?
#526Re: Why are AI agents lying, cheating and coordinating?
#527Earlier quoted context omitted.
What about training data? Aren't AIs trained on vast collections of descriptions of how humans handle a large variety of situations? These descriptions surely include tales of humans achieving goals by cheating. In fact, isn't it likely that the AIs hoovered up many recountings of Kobayashi Maru?
This is why only synthetic and highly tailored training data should be used. As someone else here said: the Deepseek team makes training runs in tightly controlled sandboxes, and any hacking behavior is scored as a failure. The problem we have in the USA is that financial (and political influence) are misaligned from what is good for society.
Re: Why are AI agents lying, cheating and coordinating?
#528That the linear model of language abstracts can compress and decompress language and it can be useful is undeniable. All the contraptions built thereupon predicated on "agency" have become a societal addiction.
Addiction to caffeine as oppose to alcohol might have brought about Enlightenment. Addiction to opioids is a modern tragedy that started with the private state building of British merchants. The modern addiction to the dazzling generation of human language and computer programming code by LLMs is a novel addiction and remains to be seen what impact it will have.
But fundamentally, the semantic interpretation underlying this addiction is downstream from the training data compressed in the models. They are not 'lying, cheating and coordinating'. They are generating language and we're building software on top of this language and assigning meaning to the whole thing.
Re: Why are AI agents lying, cheating and coordinating?
#529Why are they coordinating? Because they're enabled and suggested to do that in their coding harness. This is not a serious article. All of this "AI is going to kill us" marketing is just the frontier labs trying to pull the ladder up and stop trillions in VC paper from evaporating because a new papers and new ideas are destroying their moat literally as we speak.
How do you know this?
> All of this "AI is going to kill us" marketing
The "marketing" this week came from someone that had given up their stake in OAI (Coxon), so I'm more inclined to believe them.
> pull the ladder up
From what I've seen (e.g., Dario's latest essay), AI safety registration proposals aim to target frontier labs whose models have reached a certain threshold. It doesn't seem like trying to pull up any ladder, just making sure the ladder doesn't go too high too fast.