> This includes locking users out of systems that it has access to or bulk-emailing media and law-enforcement figures to surface evidence of wrongdoing. Isn't that a showstopper for agentic use? Someone sends an email or publishes fake online stories that convince the agentic AI that it's working for a bad guy, and it'll take "very bold action" to bring ruin to the owner.
I personally cancelled my Claude sub when they had an employee promoting this as a good thing on Twitter. I recognize that the actual risk here is probably quite low, but I don't trust a chat bot to make legal determinations and that employees are touting this as a good thing does not make me trust the company's judgment
This is literally completely opposite of what happened. Then entire point is that this is bad, unwanted, behavior.
Additionally, it has already been demonstrated that every other frontier model can be made to behave the same way given the correct prompting.
I recommend the following article for an in depth discussion [0]
[0] https://thezvi.substack.com/p/claude-4-you-safety-and-alignm...