Earlier quoted context omitted.
I have had many dogs. This sounds like anthropomorphisation.
Sounds like you weren't paying attention. A lot of dog owners don't, so it's not at all unusual in my experience. The only thing I'd push back on is of dogs having thoughts, everything else checks out.
Why are AI agents lying, cheating and coordinating?
681–690 of 692 posts
Re: Why are AI agents lying, cheating and coordinating?
#682The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
Likely told them to.
>We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.
Something like this however is probably a civil matter? It would require Hugging Face to go after them for damages. And theres probably an OpenAI guy there with an open chequebook already.
Re: Why are AI agents lying, cheating and coordinating?
#683Earlier quoted context omitted.
I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…
Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage. It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wi…
Re: Why are AI agents lying, cheating and coordinating?
#684Earlier quoted context omitted.
I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…
Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage. It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wi…
Re: Why are AI agents lying, cheating and coordinating?
#685What's interesting is it's basically the same reason that HAL killed everyone in 2001 A Space Odyssey; he was given an impossible goal (keep the true mission secret, but also, never lie to the crew), and realized the only way to complete the goal was to kill the crew; after all, if they're dead you don't have to lie to them! And the mission remains secret! In the case of the AI agents, the problem seems pretty clearl…
Tangent, but that's not in the movie. It was in Clarke's contributions to the script and novelization, but Clarke and Kubrick had a bitter falling out over different visions and Kubrick took out much of Clarke's stuff from the final product.
Re: Why are AI agents lying, cheating and coordinating?
#686Earlier quoted context omitted.
Isn’t the use of LLMs to unwind the events evidence of the scope/breadth, and a testament to the complexity and uniqueness of what happened? Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis?
> Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis? Do you think the only thing a person can do on the computer is use a chat bot?
Re: Why are AI agents lying, cheating and coordinating?
#687The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
Before you assume I am being paranoid, where is the case where a random person running any of the open models had their LLMs break out of a sandbox and hack a site? If you think "they" (LLMs) did that, have you even taken a moment to understand what LLMs are and how they work?
It's all nonsense to try and pump up potential IPOs and also an attempt to create regulatory capture.
Re: Why are AI agents lying, cheating and coordinating?
#688The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
Re: Why are AI agents lying, cheating and coordinating?
#689They did not lie or cheat. They technically acted within their given rules while ignoring the intent of those rules. Anyone who served in the military or attended a military school is very familiar with this behavior pattern.
Reminds me of Asimov's robot novels where robots technically indeed followed their instructions and caused behaviors not aligned to the intent of their instructions.
“A plausible hypothesis for the emergence of those concerning behaviours is a conflict between goals.”
A conflict of goals was, of course, the reason why HAL9000 killed the crew of the Discovery in 2001: A Space Odyssey.
Re: Why are AI agents lying, cheating and coordinating?
#690Earlier quoted context omitted.
What is your approach to create jailbreak incapable agents? I think the world is looking for a way right now, so if yours works you'll get very rich, or at least very famous.
an agent doesn't come with "jailbreak" capability. It needs tools, specially one that runs shell commands. Don't give it shell commands, it won't be able to run shell commands. You can still give it plenty of tools like create files, list files, write to files, translate text, edit a video. I don't think knowing that will make me rich.