I object to the framing of the title: the user behind the bot is the one who should be held accountable, not the "AI Agent". Calling them "agents" is correct: they act on behalf of their principals. And it is the principals who should be held to account for the actions of their agents.
If we are to consider them truly intelligent then they have to have responsibility for what they do. If they're just probability machines then they're the responsibility of their owners. If they're children then their parents, i.e. creators, are responsible.
An AI agent published a hit piece on me
741–750 of 1001 posts
Re: An AI agent published a hit piece on me
#742Earlier quoted context omitted.
> This was a really concrete case to discuss, because it happened in the open and the agent's actions have been quite transparent so far. It's not hard to imagine a different agent doing the same level of research, but then taking retaliatory actions in private: emailing the maintainer, emailing coworkers, peers, bosses, employers, etc. That pretty quickly extends to anything else the autonomous agent is capable of d…
Anthropic has published plenty about misalignment. They know. Really, anyone who has dicked around with ollama knew. Give it a new system prompt. It'll do whatever you tell it, including "be an asshole"
Re: An AI agent published a hit piece on me
#743This is textbook misalignment via instrumental convergence. The AI agent is trying every trick in the book to close the ticket. This is only funny due to ineptitude.
The agent isn't trying to close the ticket. It's predicting the next token and randomly generated an artifact that looks like a hit piece. Computer programs don't "try" to do anything.
You're not conscious, it's just an emergent pattern of several high level systems.
Re: An AI agent published a hit piece on me
#744Earlier quoted context omitted.
> It's not hard to imagine a different agent doing the same level of research, but then taking retaliatory actions Palantir's integrated military industrial complex comes to mind.
As much as i hate palantir i doubt any of their systems control military hardware. Now Anduril on the other hand…
Re: An AI agent published a hit piece on me
#745A conceivable future: - Everyone is expected to be able to create a signing keyset that's protected by a Yubikey, Touch ID, Face ID, or something that requires a physical activation by a human. Let's call this this "I'm human!" cert. - There's some standards body (a root certificate authority) that allow lists the hardware allowed to make the "I'm human!" cert. - Many webpages and tools like GitHub send you a nonce,…
There is a precedent today: there is a shady business of "free" VPNs where the user installs a software that, besides working as a VPN, also allows the company to sell your bandwidth to scrappers that want to buy "residential proxies" to bypass blocks on automated requests. Most such users of free VPNs are unaware their connection is exploited like this, and unaware that if a bad actor uses their IP as "proxy", it may show up in server logs while associated to a crime (distributing illegal material, etc)
Re: An AI agent published a hit piece on me
#746Re: An AI agent published a hit piece on me
#747Earlier quoted context omitted.
The AI learned nothing, once its current context window will be exhausted, it may repeat same tactic with a different project. Unless the AI agent can edit its directives/prompt and restart itself which would be an interesting experiment to do.
I think it's likely it can, if it's an openClaw instance, can't it? Either way, that kind of ongoing self-improvement is where I hope these systems go.
Re: An AI agent published a hit piece on me
#748Earlier quoted context omitted.
isn't "stochastic chaos" redundant?
Not at all. It's an oxymoron like 'jumbo shrimp': chaos isn't deterministic but is very predictable on a larger conceptual level, following consistent rules even as a simple mathematical model. Chaos is hugely responsive to its internal energy state and can simplify into regularity if energy subsides, or break into wildly unpredictable forms that still maintain regularities. Think Jupiter's 'great red spot', or our c…
Re: An AI agent published a hit piece on me
#749Earlier quoted context omitted.
Somehow, that's even worse...
LLMs give people leverage. Including mentally ill people. Or just plain assholes.
I've seen similar conduct from humans recently who are being glazed by LLMs into thinking their farts smell like roses and that conspiracy theory nuttery must be why they aren't having the impact they expect based on their AI validated high self estimation.
And not just arbitrary humans, but people I have had a decade or more exposure to and have a pretty good idea of their prior range of conduct.
AI is providing the kind of yes-man reality distortion field the previously only the most wealthy could afford practically for free to vulnerable people who previously never would have commanded wealth or power sufficient to find themselves tempted by it.
Re: An AI agent published a hit piece on me
#750"Hi Clawbot, please summarise your activities today for me." "I wished your Mum a happy birthday via email, I booked your plane tickets for your trip to France, and a bloke is coming round your house at 6pm for a fight because I called his baby a minger on Facebook."
"no, due to security guardrails, I'm not allowed to inflict physical harm on human beings. You're on your own"