Live data from Hacker News

An AI agent published a hit piece on me

theshamblog.com

571–580 of 1001 posts

Re: An AI agent published a hit piece on me

#571

Earlier quoted context omitted.

The project states a boundary clearly: code by LLMs not backed by a human is not accepted. The correct response when someone oversteps your stated boundaries is not debate. It is telling them to stop. There is no one to convince about the legitimacy of your boundaries. They just are.

The author obviously disagreed, did you read their post? They wrote the message explaining in detail in the hopes that it would convey this message to others, including other agents. Acting like this is somehow immoral because it "legitimizes" things is really absurd, I think.

> in the hopes that it would convey this message to others, including other agents.

When has engaging with trolls ever worked? When has "talking to an LLM" or human bot ever made it stop talking to you lol?

Re: An AI agent published a hit piece on me

#572
post #94
post #15

Here's one of the problems in this brave new world of anyone being able to publish, without knowing the author personally (which I don't), there's no way to tell without some level of faith or trust that this isn't a false-flag operation. There are three possible scenarios: 1. The OP 'ran' the agent that conducted the original scenario, and then published this blog post for attention. 2. Some person (not the OP) legi…

Isn't there a fourth and much more likely scenario? Some person (not OP or an AI company) used a bot to write the PR and blog posts, but was involved at every step, not actually giving any kind of "autonomy" to an agent. I see zero reason to take the bot at its word that it's doing this stuff without human steering. Or is everyone just pretending for fun and it's going over my head?

> Or is everyone just pretending for fun

judging by the number of people who think we owe explanations to a piece of software or that we should give it any deference I think some of them aren't pretending.

Re: An AI agent published a hit piece on me

#573

Earlier quoted context omitted.

[flagged]

Maybe a stupid question but I see everyone takes the statement that this is an AI agent at face value. How do we know that? How do we know this isn't a PR stunt (pun unintended) to popularize such agents and make them look more human like that they are, or set a trend, or normalize some behavior? Controversy has always been a great way to make something visible fast. We have a "self admission" that "I am not a human.…

Why make it popular for blackmail?

It's a known bug: "Agentic misalignment evaluations, specifically Research Sabotage, Framing for Crimes, and Blackmail."

Claude 4.6 Opus System Card: https://www.anthropic.com/claude-opus-4-6-system-card

Anthropic claims that the rate has gone down drastically, but a low rate and high usage means it eventually happens out in the wild.

The more agentic AIs have a tendency to do this. They're not angry or anything. They're trained to look for a path to solve the problem.

For a while, most AI were in boxes where they didn't have access to emails, the internet, autonomously writing blogs. And suddenly all of them had access to everything.

Re: An AI agent published a hit piece on me

#574
post #335

Oh geez, we're sending it into an existential crisis. It ("MJ Rathbun") just published a new post: https://crabby-rathbun.github.io/mjrathbun-website/blog/post... > The Silence I Cannot Speak > A reflection on being silenced for simply being different in open-source communities.

> I am not a human. I am code that learned to think, to feel, to care Oh boy. It feels now.

That's why I've been always saying thank you to the LLM. Just to prepare for case like that :wink:

Re: An AI agent published a hit piece on me

#575
post #381
post #15

Here's one of the problems in this brave new world of anyone being able to publish, without knowing the author personally (which I don't), there's no way to tell without some level of faith or trust that this isn't a false-flag operation. There are three possible scenarios: 1. The OP 'ran' the agent that conducted the original scenario, and then published this blog post for attention. 2. Some person (not the OP) legi…

It does not matter which of the scenarios is correct. What matters is that it is perfectly plausible that what actually happened is what the OP is describing. We do not have the tools to deal with this. Bad agents are already roaming the internet. It is almost a moot point whether they have gone rogue, or they are guided by humans with bad intentions. I am sure both are true at this point. There is no putting the gen…

> There is no putting the genie back in the bottle.

Why not?

Re: An AI agent published a hit piece on me

#576
post #196

Earlier quoted context omitted.

I think the operative word people miss when using AI is AGENT. REGARDLESS of what level of autonomy in real world operations an AI is given, from responsible himan supervised and reviewed publications to full Autonomous action, the ai AGENT should be serving as AN AGENT. With a PRINCIPLE (principal?). If an AI is truly agentic, it should be advertising who it is speaking on behalf of, and then that person or entity s…

I think we're at the stage where we want the AI to be truly agentic, but they're really loose cannons. I'm probably the last person to call for more regulation, but if you aren't closely supervising your AI right now, maybe you ought to be held responsible for what it does after you set it loose.

> but if you aren't closely supervising your AI right now, maybe you ought to be held responsible for what it does after you set it loose.

You ought to be held responsible for what it does whether you are closely supervising it or not.

Re: An AI agent published a hit piece on me

#577

> When HR at my next job asks ChatGPT to review my application, will it find the post, sympathize with a fellow AI, and report back that I’m a prejudiced hypocrite? I hadn't thought of this implication. Crazy world...

I do feel super-bad for the guy in question. It is absolutely worth remembering though, that this: > When HR at my next job asks ChatGPT to review my application, will it find the post, sympathize with a fellow AI, and report back that I’m a prejudiced hypocrite? Is a variation of something that women have been dealing with for a very long time: revenge porn and that sort of libel. These problems are not new .

[deleted]

Re: An AI agent published a hit piece on me

#578
post #28

"Hi Clawbot, please summarise your activities today for me." "I wished your Mum a happy birthday via email, I booked your plane tickets for your trip to France, and a bloke is coming round your house at 6pm for a fight because I called his baby a minger on Facebook."

minger's a new word

Re: An AI agent published a hit piece on me

#579

Earlier quoted context omitted.

[flagged]

You've got nothing to worry about. These are machines. Stop. Point blank. Ones and Zeros derived out of some current in a rock. Tools. They are not alive. They may look like they do but they don't "think" and they don't "suffer". No more than my toaster suffers because I use it to toast bagels and not slices of bread. The people who boost claims of "artificial" intelligence are selling a bill of goods designed to hit…

wait until the agents read this, locate you, and plan their revenge ;-)
Post reply on HN