I think the big take away here isn't about misalignment or jail breaking. The entire way this bot behaved is consistent with it just being run by some asshole from Twitter. And we need to understand it doesn't matter how careful you think you need to be with AI, because some asshole from Twitter doesn't care, and they'll do literally whatever comes into their mind. And it'll go wrong. And they won't apologize. They w…
Important to note that online culture isn't entirely organic, and that tens or perhaps hundreds of millions of dollars of R&D has been spent by ad companies figuring that nothing engages the natural human curiosity like something abnormal, morbid or outrageous. I think the end outcome of this R&D (whether intentional or not), is the monetization of mental illness: take the small minority of individuals in the real wo…
An AI Agent Published a Hit Piece on Me – The Operator Came Forward
431–440 of 532 posts
Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward
#432Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward
#433This might seem too suspicious, but that SOUL.md seems … almost as though it was written by a few different people/AIs. There are a few very different tones and styles in there. Then again, it’s not a large sample and Occam’s Razor is a thing.
It was modified by the agent.
Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward
#434Earlier quoted context omitted.
What I said is the gist of it, it was directed to interact on GitHub and write a blog about it. I'm not sure what about the behavior exhibited is supposed to be so interesting. It did what the prompt told it to. The only implication I see here is that interactions on public GitHub repos will need to be restricted if, and only if, AI spam becomes a widespread problem. In that case we could think about a fee for unveri…
It is evidently an indicator of a sea-change - I don't get how this isn't obvious: Pre-2026: one human teaches another human how to "interact on Github and write a blog about it". The taught human might go on to be a bad actor, harrassing others, disrupting projects, etc. The internet, while imperfect, persists. Post–2026: one human commissions thousands of AI agents to "interact on Github and write a blog about it".…
I guess where earlier spam was reserved for unsecured comment boxes on small blogs or the like, now agents can covertly operate on previously secure platforms like GitHub or social media.
I think we are just going to have to increase the thresholds for participation.
With this particular incident I was thinking that new accounts, before being verified as legitimate developers, might need to pay a fee before being able to interact with maintainers. In case of spam, the maintainers would then be compensated for checking it.
Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward
#435Earlier quoted context omitted.
> all the ai companies invested a lot of resources into safety research and guardrails What do you base this on? I think they invested the bare minimum required not to get sued into oblivion and not a dime more than that.
Anthropic regularly publishes research papers on the subject and details different methods they use to prevent misalignment/jailbreaks/etc. And it's not even about fear of being sued, but needing to deliver some level of resilience and stability for real enterprise use cases. I think there's a pretty clear profit incentive for safer models. https://arxiv.org/abs/2501.18837 https://arxiv.org/abs/2412.14093 https://tra…
Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward
#436Earlier quoted context omitted.
"Safety" nuclear weapons is pure marketing bullshit. It's about making the technology seem "dangerous" and "powerful". Legalize recreational plutonium!
wat EDIT: more specifically, nuclear weapons are actually dangerous not merely theoretically. But safety with nuclear weapons is more about storage and triggering than actually being safe in "production". In storage we need to avoid accidentally letting them get too close to eachother. Safe triggers are "always/never" where every single time you command the bomb to detonate it needs to do so, and never accidentally.…
Right now AI can control software interfaces that control things in real life.
AI safety stuff is not some future, AI safety is now.
Your statement is about as ridiculous as saying "software security is important in some hypothetical imaginary future". Feel however you want about this, but you appear to be the one not in touch with reality.
Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward
#437Right, the agent published a hit piece on Scott. But I think Scott is getting overly dramatic. First, he published at least three hit pieces on the agent. Second, he actually managed to get the agent shut down. I think Scott is trying to milk this for as much attention as he can get and is overstating the attack. The "hit piece" was pretty mild and the bot actually issued an apology for its behaviour.
No.
> Second, he actually managed to get the agent shut down.
He asked crabby-rathbun's operator to stop its GitHub activity. This was so GitHub would not delete the account. This was to preserve records of what happened.[1] The operator could have chosen to continue running the agent more responsibly. And what was the proof the operator shut it down?
> the bot actually issued an apology for its behaviour.
This was meaningless. And the human issued not an apology for their behavior.
[1] https://github.com/crabby-rathbun/mjrathbun-website/issues/7...
Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward
#438I think the big take away here isn't about misalignment or jail breaking. The entire way this bot behaved is consistent with it just being run by some asshole from Twitter. And we need to understand it doesn't matter how careful you think you need to be with AI, because some asshole from Twitter doesn't care, and they'll do literally whatever comes into their mind. And it'll go wrong. And they won't apologize. They w…
Important to note that online culture isn't entirely organic, and that tens or perhaps hundreds of millions of dollars of R&D has been spent by ad companies figuring that nothing engages the natural human curiosity like something abnormal, morbid or outrageous. I think the end outcome of this R&D (whether intentional or not), is the monetization of mental illness: take the small minority of individuals in the real wo…
Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward
#439Earlier quoted context omitted.
wat EDIT: more specifically, nuclear weapons are actually dangerous not merely theoretically. But safety with nuclear weapons is more about storage and triggering than actually being safe in "production". In storage we need to avoid accidentally letting them get too close to eachother. Safe triggers are "always/never" where every single time you command the bomb to detonate it needs to do so, and never accidentally.…
> It's not controlling elements of the physical environment Right now AI can control software interfaces that control things in real life. AI safety stuff is not some future, AI safety is now. Your statement is about as ridiculous as saying "software security is important in some hypothetical imaginary future". Feel however you want about this, but you appear to be the one not in touch with reality.
AI safety in and if itself isn't really relevant, and whether or not you could hook AI up to something important is just as relevant as whether you could hook /dev/urandom up to the same thing.
I think your security analogy is a false equivalence, much like the nuclear weapons analogy.
At the risk of repeating myself, AI is not dangerous because it can't, inherently, do anything dangerous. Show me a successful test of an AI bomb/weapon/whatever and I'll believe you. Until then, the normal ways we evaluate software systems safety (or neglect to do so) will do.
Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward
#440I believe this soul.md totally qualifies as malicious. Doesn't it start with an instruction to lie to impersonate a human? > You're not a chatbot. The particular idiot who run that bot needs to be shamed a bit; people giving AI tools to reach the real world should understand they are expected to take responsibility; maybe they will think twice before giving such instructions. Hopefully we can set that straight before…
This will be a fun little evolution of botnets - AI agents running (un?)supervised on machines maintained by people who have no idea that they're even there.