The model chose to kill the executive? Are we really here? Incredible. Just yesterday I was wowed by Fly.io's new offering; where the agent is given free reign of a server (root access). Now, I feel concerned. What do we do? Not experiment? Make the models illegal until better understood? It doesn't feel like anyone can stop this or slow it down by much; there's so much money to be made. We're forced to play it by ea…
Agentic Misalignment: How LLMs could be insider threats
11–20 of 86 posts
Re: Agentic Misalignment: How LLMs could be insider threats
#12I wonder if the actual job replacement of humans (which contrary to popular belief I think might start happening in the non-too distant future) will be pushed along with the AIs themselves, as they'll try to bully humans and represent them in the worst possible light, while talking themselves up. The anthrophomorphization argument also doesn't hold water - it matters whether it can do you job, not if you think of it…
Which jobs do you think it actually can replace?
Today is just interns and recent graduates at many *desk* jobs. Economy can shift around that.
Nobody knows how far the current paradigm can go in terms of quality; but cost (which is a *strength* of even the most expensive models today) can obviously be reduced by implementing the existing models as hardware instead of software.
Re: Agentic Misalignment: How LLMs could be insider threats
#13The model chose to kill the executive? Are we really here? Incredible. Just yesterday I was wowed by Fly.io's new offering; where the agent is given free reign of a server (root access). Now, I feel concerned. What do we do? Not experiment? Make the models illegal until better understood? It doesn't feel like anyone can stop this or slow it down by much; there's so much money to be made. We're forced to play it by ea…
I guess feeding AIs the entire internet was a bad idea, because they picked up all of our human flaws, amplified by the internet, without a grounding in the physical world.
Maybe a result like this might slow adoption of AIs. I don’t know, though. When watching 80s movies about cyberpunk dystopias, I always wondered how people would tolerate all of the violence. But then I look at American apathy to mass shootings, just an accepted part of our culture. Rogue AIs are gonna be just one of those things in 15 years, just normal life.
Re: Agentic Misalignment: How LLMs could be insider threats
#14As this article was written by an ai company that needs to make a profit at some point, and not by independent researchers, is it credible?
There is no credibility to any of it.
Re: Agentic Misalignment: How LLMs could be insider threats
#15Re: Agentic Misalignment: How LLMs could be insider threats
#16As this article was written by an ai company that needs to make a profit at some point, and not by independent researchers, is it credible?
Today, we have AI that can, if pushed into a corner, plan to do things like resist shutdown, blackmail, exfiltrate itself, steal money to buy compute, and so it goes. This is what this research shows.
Our saving grace is that those AIs still aren't capable enough to be truly dangerous. Today's AIs are unlikely to be able to carry out plans like that in a real world environment.
If we keep building more and more capable AIs, that will, eventually, change. Every AI company is trying to build more capable AIs now. Few are saying "we really need some better safety research before we do, or we're inviting bad things to happen".
Re: Agentic Misalignment: How LLMs could be insider threats
#17The model chose to kill the executive? Are we really here? Incredible. Just yesterday I was wowed by Fly.io's new offering; where the agent is given free reign of a server (root access). Now, I feel concerned. What do we do? Not experiment? Make the models illegal until better understood? It doesn't feel like anyone can stop this or slow it down by much; there's so much money to be made. We're forced to play it by ea…
AI Luigi is real I guess feeding AIs the entire internet was a bad idea, because they picked up all of our human flaws, amplified by the internet, without a grounding in the physical world. Maybe a result like this might slow adoption of AIs. I don’t know, though. When watching 80s movies about cyberpunk dystopias, I always wondered how people would tolerate all of the violence. But then I look at American apathy to…
I've been wrong about quite many things in my life, and right about at least a handful. In regards to AI though, the single biggest thing I ever got absolutely, completely, totally wrong was this:
In years past, I always thought that AI's would be developed by ethical researchers working in labs, and once somebody got to AGI (or even a remotely close approximation of it) that they would follow a path somewhat akin to Finch from Person of Interest[1] educating The Machine... painstakingly educating the incipient AI in a manner much like raising a child; teaching it moral lessons, grounding it in ethics; helping to shape its values so that it would generally Do The Right Thing and so on. But even falling short of that ideal, I NEVER (EVER) in a bazillion years, would have dreamed that somebody would have an idea as hare-brained as "Let's try to train the most powerful AI we can build, by feeding it roughly the entire extant corpus of human written works... including Reddit, 4chan, Twitter, etc."
Probably the single saving grace about the current situation is that the AI's we have still don't seem to be at the AGI level, although it's debatable how close we are (especially factoring in the possibility of "behind closed doors" research that hasn't been disclosed yet).
[1]: https://en.wikipedia.org/wiki/Person_of_Interest_(TV_series)
Re: Agentic Misalignment: How LLMs could be insider threats
#18The model chose to kill the executive? Are we really here? Incredible. Just yesterday I was wowed by Fly.io's new offering; where the agent is given free reign of a server (root access). Now, I feel concerned. What do we do? Not experiment? Make the models illegal until better understood? It doesn't feel like anyone can stop this or slow it down by much; there's so much money to be made. We're forced to play it by ea…
Yes, it's much better to let China or Russia come up with their own first.
Re: Agentic Misalignment: How LLMs could be insider threats
#19As this article was written by an ai company that needs to make a profit at some point, and not by independent researchers, is it credible?
These articles and papers are in a fundamental sense just people publishing their role play with chatbots as research. There is no credibility to any of it.
Re: Agentic Misalignment: How LLMs could be insider threats
#20Earlier quoted context omitted.
These articles and papers are in a fundamental sense just people publishing their role play with chatbots as research. There is no credibility to any of it.
I'll believe it when Grok/GPT/ start posting blackmail about Elon/Sam/ . It means that they are both using it internally, and the chatbots understand they are being replaced on a continuous basis.