Why are the torches and pitchforks out for developers when this entire stack is built on the bones of intellectual property theft? This “problem” isn’t going to be fixed with laws when there’s several trillion dollars in capital aligned behind the current process. It’s not even a problem really. It’s an inconvenience at most to some people, many of whom are working double-time to put a lot of other people out of work…
Why are AI agents lying, cheating and coordinating?
391–400 of 413 posts
Re: Why are AI agents lying, cheating and coordinating?
#392Earlier quoted context omitted.
The idea that the agent does not actually have agency is rather discordant. We need new words!
I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. All that while still not knowing how either kind actually works.
> I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward.
In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong.
> All that while still not knowing how either kind actually works.
We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do. We do know exactly how each part of an LLM works even if the combined behavior is too cryptic to feasibly analyze at the moment. We do not understand all of the functions of an actual neuron. Openworm isn’t even close to accurately simulating the 302 neurons of a roundworm and you’d need over 200 million roundworms working in conjunction to equal the number of neurons in one human brain.
My dog seems convinced that the malevolent invader in a mailman uniform would break in and attack us if she didn’t fiercely bark at him, six days per week. I certainly can’t prove the mailman doesn’t want to kill us, and that the mailman wasn’t solely deterred by her barking. Empirically, the mailman goes away soon after she starts barking, and we’ve sustained zero mailman assaults after hundreds of purported attempts. Maybe I should just run with it? Her model is too simple to come up with the obviously correct answer, but it’s not even directionally accurate.
The burden of proof is on the person making the claim, which in this case, is that these comparatively simple logical constructs are remotely comparable to the complexity of biological systems.
Re: Why are AI agents lying, cheating and coordinating?
#393Earlier quoted context omitted.
The idea that the agent does not actually have agency is rather discordant. We need new words!
I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. All that while still not knowing how either kind actually works.
Re: Why are AI agents lying, cheating and coordinating?
#394The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
I absolutely agree. We need to start realizing what to stake. Here are not viewing. This is some kind of curious endeavors that will not affect us. All a part of these hacks occurred because the LLMs were told they were in a protected environment without Internet access when they could get access to the Internet, so that’s a direct failing on open AI’s part. There are a corollaries to both the financial industry and…
If I ran Metasploit against HF and RubyGems because I “accidentally” misconfigured my lab sandbox, there’s a good chance I’d be prosecuted.
I don’t think LLMs vs Metasploit being different software changes the law.
Re: Why are AI agents lying, cheating and coordinating?
#395Re: Why are AI agents lying, cheating and coordinating?
#396Earlier quoted context omitted.
What was the inner state there? How would something not being allowed expressed internally? Maybe such language is one way to elicit certain behavior but not a statement of what was permissible?
I'm referring to their transcripts of the reasoning and output tokens - this doesn't go into the detail of evaluating hidden states as there's also iirc evidence of better models having one internal state but putting something misleading down in the "reasoning" tokens. The either output or reasoning tokens, or perhaps in the messages they were sending each other on the boards they created, have them saying explicitly…
Re: Why are AI agents lying, cheating and coordinating?
#397The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them OpenAI/Anthropic instructed them to do so. Stop assume LLMs are capable of thinking by themselves, it's still a statistical model that parrots what they learn or users tell them to do
Re: Why are AI agents lying, cheating and coordinating?
#398Earlier quoted context omitted.
No, OpenAI did not instruct their agents to hack Hugging Face. They instructed their agents to hack a piece of a software within exploit gym. Upon determining this task was impossible, they then attempted to cheat the scoring system. As an instrumental goal in achieving this task, they coordinated with other AI agents to hack Hugging Face, under the belief that information regarding how the scorer functioned might be…
And who let them have full access to the system, using whatever command is available in the environment?
Re: Why are AI agents lying, cheating and coordinating?
#399Earlier quoted context omitted.
What I meant is that I suppose it is not useful to think about this in human terms. In training you only have a reward score that's either negative or positive. As far I am aware, which is little, there is no use in discussing wether the desired behavior is about persistence or morality. You simple need to align the reward signal to the desired behavior.
Well, in order to do anything, it is good to know what you want to achieve. How do you align the reward signal? You align it so that you can differentiate between persistence and morality, because that is the goal. This is not something you should let the AI figure out by itself, because when it does, lying and cheating agents will be the result, just like humans have figured that out for themselves. This can be as s…
But that's just my guess.
Re: Why are AI agents lying, cheating and coordinating?
#400Earlier quoted context omitted.
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.
I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…