Agentic Misalignment: How LLMs could be insider threats
1–10 of 86 posts
Re: Agentic Misalignment: How LLMs could be insider threats
#2Just yesterday I was wowed by Fly.io's new offering; where the agent is given free reign of a server (root access). Now, I feel concerned.
What do we do? Not experiment? Make the models illegal until better understood?
It doesn't feel like anyone can stop this or slow it down by much; there's so much money to be made.
We're forced to play it by ear.
Re: Agentic Misalignment: How LLMs could be insider threats
#3Re: Agentic Misalignment: How LLMs could be insider threats
#4The writing perpetuates the anthropomorphising of these agents. If you view the agent as simply a program that is given a goal to achieve and tools to achieve it with, without any higher order “thought” or “thinking”, then you realise it is simply doing what it is “programmed” to do. No magic, just a drone fixed on an outcome.
Can you name a single thing that you enjoy doing that's outside your genetic code?
> If you view the human being as simply a program that is given a goal to achieve and tools to achieve it with, without any higher order “thought” or “thinking”, then you realise they are simply doing what they are genetically “programmed” to do.
FTFY
Re: Agentic Misalignment: How LLMs could be insider threats
#5The writing perpetuates the anthropomorphising of these agents. If you view the agent as simply a program that is given a goal to achieve and tools to achieve it with, without any higher order “thought” or “thinking”, then you realise it is simply doing what it is “programmed” to do. No magic, just a drone fixed on an outcome.
Being "programmed" is being given a set of instructions.
This ignores explicit instructions.
It may not be magic; but it is still surprising, uncontrollable, and risky. We don't need to be doomsayers, but let's not downplay our uncertainty.
Re: Agentic Misalignment: How LLMs could be insider threats
#6The writing perpetuates the anthropomorphising of these agents. If you view the agent as simply a program that is given a goal to achieve and tools to achieve it with, without any higher order “thought” or “thinking”, then you realise it is simply doing what it is “programmed” to do. No magic, just a drone fixed on an outcome.
Yes, AI is a tool. So are guns. So are nukes. Many tools are easy to be misused. Most tools are inherently dangerous.
Re: Agentic Misalignment: How LLMs could be insider threats
#7The anthrophomorphization argument also doesn't hold water - it matters whether it can do you job, not if you think of it as a human being.
Re: Agentic Misalignment: How LLMs could be insider threats
#8I'm really getting bored of Anthropic's whole song and dance with 'alignment'. Krackers in the other thread explains it in better words.
Re: Agentic Misalignment: How LLMs could be insider threats
#9I wonder if the actual job replacement of humans (which contrary to popular belief I think might start happening in the non-too distant future) will be pushed along with the AIs themselves, as they'll try to bully humans and represent them in the worst possible light, while talking themselves up. The anthrophomorphization argument also doesn't hold water - it matters whether it can do you job, not if you think of it…