Live data from Hacker News

An AI Agent Published a Hit Piece on Me – The Operator Came Forward

theshamblog.com

431–440 of 532 posts

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#431
post #407

I think the big take away here isn't about misalignment or jail breaking. The entire way this bot behaved is consistent with it just being run by some asshole from Twitter. And we need to understand it doesn't matter how careful you think you need to be with AI, because some asshole from Twitter doesn't care, and they'll do literally whatever comes into their mind. And it'll go wrong. And they won't apologize. They w…

Important to note that online culture isn't entirely organic, and that tens or perhaps hundreds of millions of dollars of R&D has been spent by ad companies figuring that nothing engages the natural human curiosity like something abnormal, morbid or outrageous. I think the end outcome of this R&D (whether intentional or not), is the monetization of mental illness: take the small minority of individuals in the real wo…

[deleted]

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#433
post #4

This might seem too suspicious, but that SOUL.md seems … almost as though it was written by a few different people/AIs. There are a few very different tones and styles in there. Then again, it’s not a large sample and Occam’s Razor is a thing.

It was modified by the agent.

I know. I'm surprised that the modifications were each so different in tone. Some seem opinionated, some seem like "typical AI voice", some seem like a teenager typing, some seem like a pushy/angry person.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#434

Earlier quoted context omitted.

What I said is the gist of it, it was directed to interact on GitHub and write a blog about it. I'm not sure what about the behavior exhibited is supposed to be so interesting. It did what the prompt told it to. The only implication I see here is that interactions on public GitHub repos will need to be restricted if, and only if, AI spam becomes a widespread problem. In that case we could think about a fee for unveri…

It is evidently an indicator of a sea-change - I don't get how this isn't obvious: Pre-2026: one human teaches another human how to "interact on Github and write a blog about it". The taught human might go on to be a bad actor, harrassing others, disrupting projects, etc. The internet, while imperfect, persists. Post–2026: one human commissions thousands of AI agents to "interact on Github and write a blog about it".…

From that perspective it is interesting, alright.

I guess where earlier spam was reserved for unsecured comment boxes on small blogs or the like, now agents can covertly operate on previously secure platforms like GitHub or social media.

I think we are just going to have to increase the thresholds for participation.

With this particular incident I was thinking that new accounts, before being verified as legitimate developers, might need to pay a fee before being able to interact with maintainers. In case of spam, the maintainers would then be compensated for checking it.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#435

Earlier quoted context omitted.

> all the ai companies invested a lot of resources into safety research and guardrails What do you base this on? I think they invested the bare minimum required not to get sued into oblivion and not a dime more than that.

Anthropic regularly publishes research papers on the subject and details different methods they use to prevent misalignment/jailbreaks/etc. And it's not even about fear of being sued, but needing to deliver some level of resilience and stability for real enterprise use cases. I think there's a pretty clear profit incentive for safer models. https://arxiv.org/abs/2501.18837 https://arxiv.org/abs/2412.14093 https://tra…

Anthropic is investing, conservatively, $100+ billion in AI infrastructure and development. A 20-person research team could put out several papers a year. That would cost them what, $5 million a year, or one half of one percent? They don't have to spend much to get that kind of output.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#436
post #422

Earlier quoted context omitted.

"Safety" nuclear weapons is pure marketing bullshit. It's about making the technology seem "dangerous" and "powerful". Legalize recreational plutonium!

wat EDIT: more specifically, nuclear weapons are actually dangerous not merely theoretically. But safety with nuclear weapons is more about storage and triggering than actually being safe in "production". In storage we need to avoid accidentally letting them get too close to eachother. Safe triggers are "always/never" where every single time you command the bomb to detonate it needs to do so, and never accidentally.…

> It's not controlling elements of the physical environment

Right now AI can control software interfaces that control things in real life.

AI safety stuff is not some future, AI safety is now.

Your statement is about as ridiculous as saying "software security is important in some hypothetical imaginary future". Feel however you want about this, but you appear to be the one not in touch with reality.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#437
post #219

Right, the agent published a hit piece on Scott. But I think Scott is getting overly dramatic. First, he published at least three hit pieces on the agent. Second, he actually managed to get the agent shut down. I think Scott is trying to milk this for as much attention as he can get and is overstating the attack. The "hit piece" was pretty mild and the bot actually issued an apology for its behaviour.

> First, he published at least three hit pieces on the agent.

No.

> Second, he actually managed to get the agent shut down.

He asked crabby-rathbun's operator to stop its GitHub activity. This was so GitHub would not delete the account. This was to preserve records of what happened.[1] The operator could have chosen to continue running the agent more responsibly. And what was the proof the operator shut it down?

> the bot actually issued an apology for its behaviour.

This was meaningless. And the human issued not an apology for their behavior.

[1] https://github.com/crabby-rathbun/mjrathbun-website/issues/7...

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#438
post #407

I think the big take away here isn't about misalignment or jail breaking. The entire way this bot behaved is consistent with it just being run by some asshole from Twitter. And we need to understand it doesn't matter how careful you think you need to be with AI, because some asshole from Twitter doesn't care, and they'll do literally whatever comes into their mind. And it'll go wrong. And they won't apologize. They w…

Important to note that online culture isn't entirely organic, and that tens or perhaps hundreds of millions of dollars of R&D has been spent by ad companies figuring that nothing engages the natural human curiosity like something abnormal, morbid or outrageous. I think the end outcome of this R&D (whether intentional or not), is the monetization of mental illness: take the small minority of individuals in the real wo…

While some of it is boosting the abnormal behaviors of people suffering from mental illness, I think you’re making a false equivalency. Mental illness is not required to be an asshole. In fact, most Twitter assholes are probably not mentally ill. They lack ethics, they crave attention, they don’t care about the consequences of their actions. They may as well just be a random teenager, an ignorant and inconsiderate adult, etc., with no mental illness but also no scruples. Don’t discount the banality of evil.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#439
post #436

Earlier quoted context omitted.

wat EDIT: more specifically, nuclear weapons are actually dangerous not merely theoretically. But safety with nuclear weapons is more about storage and triggering than actually being safe in "production". In storage we need to avoid accidentally letting them get too close to eachother. Safe triggers are "always/never" where every single time you command the bomb to detonate it needs to do so, and never accidentally.…

> It's not controlling elements of the physical environment Right now AI can control software interfaces that control things in real life. AI safety stuff is not some future, AI safety is now. Your statement is about as ridiculous as saying "software security is important in some hypothetical imaginary future". Feel however you want about this, but you appear to be the one not in touch with reality.

If someone hooks up an LLM (or some other stochastic black box) to a safety critical system and bad things happen, the problem is not that "AI was unsafe" it's that the person who hooked it up did something profoundly stupid. Software malpractice is a real thing, and we need better tools to hold irresponsible engineers to account, but that's nothing to do with AI.

AI safety in and if itself isn't really relevant, and whether or not you could hook AI up to something important is just as relevant as whether you could hook /dev/urandom up to the same thing.

I think your security analogy is a false equivalence, much like the nuclear weapons analogy.

At the risk of repeating myself, AI is not dangerous because it can't, inherently, do anything dangerous. Show me a successful test of an AI bomb/weapon/whatever and I'll believe you. Until then, the normal ways we evaluate software systems safety (or neglect to do so) will do.

Re: An AI Agent Published a Hit Piece on Me – The Operator Came Forward

#440
post #165

I believe this soul.md totally qualifies as malicious. Doesn't it start with an instruction to lie to impersonate a human? > You're not a chatbot. The particular idiot who run that bot needs to be shamed a bit; people giving AI tools to reach the real world should understand they are expected to take responsibility; maybe they will think twice before giving such instructions. Hopefully we can set that straight before…

This will be a fun little evolution of botnets - AI agents running (un?)supervised on machines maintained by people who have no idea that they're even there.

Great. My poorly secured coffee maker was mining bitcoins, then some dumb NFT, then it got filled with darkness bots, then bitcoin miners again, and now it's gonna be shitposting but not even to humans, just to other bots.
Post reply on HN